5g mifi device integrating voice and visual recognition

By integrating voice and visual recognition into 5G MiFi devices, seamless collaborative processing of voice and visual information is achieved, solving the problem of low efficiency in network optimization and fault diagnosis of portable devices in dynamic environments, and improving user experience and network stability.

CN121568145BActive Publication Date: 2026-04-17HUNAN YOUPIN IOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN YOUPIN IOT TECH CO LTD
Filing Date
2026-01-26
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing portable devices struggle to achieve seamless collaborative processing of voice and visual information, especially in dynamic environments, resulting in low efficiency for network optimization and fault diagnosis in complex scenarios.

Method used

5G MiFi devices that integrate voice and visual recognition use semantic understanding algorithms to parse voice commands and combine them with visual gesture recognition to achieve multimodal fusion, dynamically adjust network connectivity and fault diagnosis, and provide real-time optimization suggestions.

Benefits of technology

In complex multi-user network scenarios, it achieves real-time and accurate network optimization and fault diagnosis, improves network security, stability and user experience, and solves the problems of network congestion, uncontrolled permissions and difficulty in fault location caused by multiple concurrent accesses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121568145B_ABST
    Figure CN121568145B_ABST
Patent Text Reader

Abstract

This invention discloses a 5G MiFi device integrating voice and visual recognition. It acquires user intent through voice semantic understanding, combines the current downlink rate and carrier aggregation status with a cloud-based large model to generate optimization suggestions, and integrates user gesture recognition results collected by a visual module for access control, ensuring that only authorized users can trigger multi-user network resource adjustments. When initial assessment indicates that multi-user access exceeds a threshold, a wireless connection expansion mechanism is automatically activated and resources are dynamically allocated, ultimately forming an optimized device connection list. Furthermore, through dual-modal fusion of voice commands and visual diagnostic data, accurate fault diagnosis and identification are achieved in complex scenarios. This invention effectively solves the three core problems caused by concurrent multi-user access: network congestion, uncontrolled access, and difficulty in fault location, significantly improving network security, stability, and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of MIFI device technology, and in particular discloses a 5G MIFI device integrating voice and visual recognition. Background Technology

[0002] In the field of modern communication technology, the combination of high-speed networks and intelligent interaction is becoming an important direction driving innovation in personal and industry applications. This integration is not only about data transmission speed, but also about how to improve user experience and scenario adaptability through technological means. With the popularization of 5G technology, devices are no longer just responsible for networking, but are gradually evolving towards intelligence, portability, and multi-functionality. Especially in scenarios such as mobile office and outdoor work, intelligent network terminal devices are particularly crucial. However, current portable network devices on the market still have significant shortcomings in functional integration.

[0003] While many devices offer high-speed network connectivity, they often struggle to effectively handle interactive needs in complex environments, particularly when processing both voice commands and visual information simultaneously. These devices lack the comprehensive processing capabilities to handle multiple information sources. This deficiency forces users to frequently switch tools or rely on additional devices, reducing efficiency and limiting the device's application potential in diverse scenarios.

[0004] A deeper challenge lies in achieving coordinated processing of voice and visual information in portable devices, a key technological hurdle. Voice interaction requires devices to accurately recognize commands and respond quickly, while visual perception demands the ability to capture environmental changes and make judgments. These two aspects are often difficult to achieve simultaneously and efficiently with the limited hardware resources of portable devices. This is especially true when devices need to handle multiple tasks in dynamic environments, such as simultaneously recognizing user voice commands and environmental gestures. The conflict between hardware computing power and information integration capabilities becomes particularly pronounced. This conflict directly impacts device performance in complex scenarios. For example, during outdoor inspections, devices may be unable to simultaneously process voice commands and visual environmental changes, preventing inspectors from obtaining timely network optimization suggestions or device status feedback.

[0005] Therefore, how to achieve seamless collaborative processing of voice and visual information in resource-constrained portable devices and ensure rapid response to user needs in dynamic scenarios has become a key issue that this research urgently needs to address. Summary of the Invention

[0006] This invention provides a 5G MiFi device that integrates voice and visual recognition, aiming to solve at least one of the defects existing in the prior art.

[0007] This invention relates to a 5G MiFi device integrating voice and visual recognition, comprising:

[0008] The semantic result acquisition module is used to acquire audio signals from user voice input, acquire audio signals through an audio acquisition device, and use a semantic understanding algorithm to parse the voice command content of the audio signals to obtain semantic results related to network parameter query or signal optimization.

[0009] The optimization suggestion feedback data acquisition module is used to determine the instruction type based on the semantic results. If the semantic results include network parameter queries, the current downlink rate and carrier aggregation status are extracted from the network communication module and uploaded to the cloud big model for processing through the end-cloud collaborative communication channel to obtain optimization suggestion feedback data.

[0010] The preliminary evaluation result acquisition module is used to integrate the visual data collected by the visual module with the optimization suggestion feedback data, and to analyze the user's gestures using a gesture recognition algorithm. If the user's gestures match the permission control requirements, the feedback data is combined with the image scene perception to obtain the preliminary evaluation result of the multi-user network access permission.

[0011] The device connection list acquisition module is used to determine the device access requirements based on the preliminary assessment results. If the preliminary assessment results show that multiple users access the device more than the threshold, the wireless connection expansion mechanism is activated to allocate access resources and obtain a dynamically adjusted device connection list.

[0012] The complex scene recognition conclusion acquisition module is used to obtain real-time signal strength indicators from the device connection list, integrate voice command content and visual diagnostic data using a dual-modal fusion algorithm, determine fault diagnosis output, and obtain complex scene recognition conclusions.

[0013] Furthermore, the semantic result acquisition module includes:

[0014] The instruction energy feature extraction unit is used to acquire the sampling sequence obtained by the audio acquisition device and convert it, and extract the instruction energy feature from the sampling sequence.

[0015] The keyword extraction unit is used to generate an intent vector by encoding the instruction energy features using a deep neural network model, and then extracts search keywords based on the intent vector.

[0016] The semantic result output unit is used to locate the parameter index or generate optimization strategies based on the search keywords, and output semantic results related to network parameter query or signal optimization based on the parameter index or optimization strategies.

[0017] Furthermore, the module for obtaining feedback data on optimization suggestions includes:

[0018] The instruction type label determination unit is used to parse the intent slot information contained in the semantic result and compare the intent slot information with the preset instruction set to determine the instruction type label.

[0019] The encrypted request message encapsulation unit is used to read the underlying registers to generate the original network status data and encapsulate it into an encrypted request message if the instruction type label is a network parameter query category.

[0020] The network performance bottleneck feature acquisition unit is used to send encrypted request messages to the cloud server, and the cloud server inputs the large model to obtain the network performance bottleneck features.

[0021] The optimization suggestion feedback data generation unit is used to generate optimization suggestion feedback data containing configuration adjustment parameters based on network performance bottleneck characteristics.

[0022] Furthermore, the module for obtaining preliminary assessment results includes:

[0023] The spatiotemporal synchronization data packet generation unit is used to receive optimization suggestion feedback data sent from the cloud and visual acquisition data collected by the vision module, and to align the optimization suggestion feedback data and visual acquisition data to generate a spatiotemporal synchronization data packet.

[0024] The user gesture command encoding acquisition unit is used to extract image frames from the spatiotemporal synchronization data packet and calculate the user gesture command encoding.

[0025] A legitimate control request signal generation unit is used to generate a legitimate control request signal if the user's gesture command encoding matches the preset permission control requirements.

[0026] The preliminary evaluation result output unit is used to construct a scene topology diagram based on legitimate control request signals, and to map the optimization suggestion feedback data to the scene topology diagram to output the preliminary evaluation results of multi-user network access permissions.

[0027] Furthermore, the device connection list acquisition module includes:

[0028] The device access demand load calculation unit is used to parse the preliminary assessment results of multi-user network access permissions and calculate the device access demand load. The device access demand load is determined based on the number of terminal requests and the service bandwidth requirement value.

[0029] The available spectrum resource block filtering unit is used to activate the wireless connection extension mechanism to filter available spectrum resource blocks if the device access demand load is greater than the single node access capacity threshold. Available spectrum resource blocks are extracted from the spare radio frequency channel.

[0030] The multi-link concurrent access configuration table generation unit is used to generate a multi-link concurrent access configuration table based on available spectrum resource blocks. The multi-link concurrent access configuration table is used to perform device registration and link migration.

[0031] The device connection list acquisition unit is used to update the mapping relationship based on the multi-link concurrent access configuration table to obtain a dynamically adjusted device connection list.

[0032] Furthermore, the complex scene recognition conclusion acquisition module includes:

[0033] The voiceprint spectrum feature conversion unit is used to obtain the real-time signal strength and voice command content in the device connection list, and convert the voice command content into voiceprint spectrum features.

[0034] The dual-modal feature acquisition unit is used to generate enhanced texture features by combining visual diagnostic data if the real-time signal strength is lower than the threshold, and to aggregate the enhanced texture features with the acoustic spectrum features to obtain dual-modal features.

[0035] The scene semantic label determination unit is used to input bimodal features into the fault classification model to obtain the fault probability distribution, and determine the scene semantic label based on the bimodal features;

[0036] The complex scene recognition conclusion acquisition unit is used to aggregate the fault probability distribution and scene semantic labels to generate a diagnostic decision vector, parse the diagnostic decision vector to determine the fault diagnosis output, and obtain the complex scene recognition conclusion.

[0037] The beneficial effects achieved by this invention are as follows:

[0038] This invention provides a 5G MiFi device integrating voice and visual recognition, capable of offering real-time and accurate solutions for user-initiated network parameter queries or signal optimization requests in complex multi-user network scenarios such as homes or small offices. The invention acquires user intent through voice semantic understanding, combines the current downlink rate and carrier aggregation status with a cloud-based large model to generate optimization suggestions, and integrates user gesture recognition results collected by a visual module for access control, ensuring only authorized users can trigger multi-user network resource adjustments. When initial assessment indicates that multi-user access exceeds a threshold, a wireless connection expansion mechanism is automatically activated and resources are dynamically allocated, ultimately forming an optimized device connection list. Through dual-modal fusion of voice commands and visual diagnostic data, accurate fault diagnosis and identification are achieved in complex scenarios. This invention effectively solves the three core problems of network congestion, uncontrolled access, and difficulty in fault location caused by concurrent multi-user access, significantly improving network security, stability, and user experience. Attached Figure Description

[0039] Figure 1 This is a functional block diagram of an embodiment of the 5G MiFi device integrating voice and visual recognition according to the present invention.

[0040] Explanation of icon numbers:

[0041] 10. Semantic result acquisition module; 20. Optimization suggestion feedback data acquisition module; 30. Preliminary evaluation result acquisition module; 40. Device connection list acquisition module; 50. Complex scene recognition conclusion acquisition module. Detailed Implementation

[0042] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0043] like Figure 1 As shown, the first embodiment of the present invention proposes a 5G MiFi device integrating voice and visual recognition, including a semantic result acquisition module 10, an optimization suggestion feedback data acquisition module 20, a preliminary evaluation result acquisition module 30, a device connection list acquisition module 40, and a complex scene recognition conclusion acquisition module 50. The semantic result acquisition module 10 is used to acquire audio signals from user voice input, acquire audio signals through an audio acquisition device, and parse the voice command content of the audio signals using a semantic understanding algorithm to obtain semantic results related to network parameter queries or signal optimization. The optimization suggestion feedback data acquisition module 20 is used to determine the command type based on the semantic results. If the semantic results include network parameter queries, it extracts the current downlink rate and carrier aggregation status from the network communication module and uploads them to the cloud for large-scale model processing through the end-to-cloud collaborative communication channel to obtain optimization suggestion feedback data. The evaluation result acquisition module 30 is used to integrate the visual data collected by the visual module with the optimization suggestion feedback data, and analyze the user's gestures using a gesture recognition algorithm. If the user's gestures match the permission control requirements, the feedback data is combined with image scene perception to obtain a preliminary evaluation result of multi-user network access permissions. The device connection list acquisition module 40 is used to determine the device access requirements based on the preliminary evaluation results. If the preliminary evaluation results show that multiple users access the device more than the threshold, the wireless connection expansion mechanism is activated to allocate access resources and obtain a dynamically adjusted device connection list. The complex scene recognition conclusion acquisition module 50 is used to obtain real-time signal strength indicators from the device connection list, integrate voice command content and visual diagnostic data using a dual-modal fusion algorithm, determine the fault diagnosis output, and obtain the complex scene recognition conclusion.

[0044] The semantic result acquisition module 10 captures the audio signal of the user's voice input through the audio acquisition device (such as a high-sensitivity microphone array) on the device. After filtering out environmental noise, it uses a semantic understanding algorithm adapted to spoken command scenarios (such as a pre-trained large speech model) to perform speech-to-text, intent recognition, and keyword extraction on the audio signal. It accurately parses out the core semantics corresponding to the command, focuses on two core needs: network parameter query or signal optimization, and outputs structured semantic results. This provides the core basis for subsequent command type determination and device control startup, and is the source link for realizing voice interaction control of 5G MIFI devices.

[0045] The optimization suggestion feedback data acquisition module 20 takes the semantic results output by the semantic result acquisition module 10 as input and determines the instruction type through preset instruction classification rules. If the semantic result is a network parameter query request, it calls the device's built-in network communication module to extract core parameters such as the downlink rate and carrier aggregation status of the current 5G network in real time. Then, it uploads the parameters to the cloud big model through the end-to-cloud collaborative communication channel. The cloud big model analyzes and calculates the data in combination with global data such as network topology and base station load to generate targeted network optimization suggestion feedback data, providing data support for subsequent permission assessment and device adjustment. If it is another instruction type, it executes the corresponding preset process, which is a key bridge connecting voice interaction and intelligent device control.

[0046] The preliminary evaluation result acquisition module 30 uses the optimization suggestion feedback data output by the optimization suggestion feedback data acquisition module 20 and the environmental visual data collected by the device vision module (such as a high-definition camera) as dual input sources. On the one hand, it uses a deep learning gesture recognition algorithm to detect and classify user gestures in the visual data and verify whether the gestures match the device's preset permission control requirements. On the other hand, it integrates and analyzes the optimization suggestion feedback data that has passed the permission verification with the scene perception information (such as the number of network users and environmental occlusion) in the visual data to comprehensively judge the rationality of the multi-user network access needs in the current scenario. Finally, it outputs the preliminary evaluation results of the multi-user network access permissions, providing a dual basis of permissions and scenario for subsequent device connection resource allocation.

[0047] Based on the preliminary evaluation results output by the preliminary evaluation result acquisition module 30, the device connection list acquisition module 40 determines whether the number of devices currently applying for access exceeds the preset connection threshold of the MIFI device. If the preliminary evaluation results show that the number of multiple users accessing the device exceeds the threshold, the device's built-in wireless connection extension mechanism is automatically activated. By adjusting channel bandwidth, optimizing frequency band allocation, and enabling load balancing algorithms, the device access resources are dynamically allocated to prioritize the connection stability of high-priority devices. If the threshold is not exceeded, the existing connection status is maintained, and a dynamically adjusted device connection list is finally generated, realizing intelligent and refined management and control of 5G MIFI device access resources.

[0048] The complex scene recognition conclusion acquisition module 50 extracts the real-time signal strength indicators of each access device from the device connection list generated by the device connection list acquisition module 40. At the same time, it adopts a dual-modal fusion algorithm to deeply fuse the voice command content parsed by the semantic result acquisition module 10 (such as network lag, disconnection and other problems reported by users) with the environmental visual diagnostic data collected by the vision module (such as device placement, signal obstruction, and surrounding interference sources). Through feature alignment, weight allocation and other operations, a multi-dimensional fault diagnosis model is constructed to accurately locate the root cause of network anomalies, output fault diagnosis results and complex scene recognition conclusions, provide users with fault resolution strategies, and feed the diagnostic data back to the device system to optimize the subsequent control process, forming a complete closed loop of "interaction-control-diagnosis-optimization".

[0049] Furthermore, the 5G MiFi device integrating voice and visual recognition provided in this embodiment includes a semantic result acquisition module 10 comprising an instruction energy feature extraction unit, a search keyword extraction unit, and a semantic result output unit. The instruction energy feature extraction unit is used to acquire the sampling sequence obtained by the audio acquisition device and extract instruction energy features from the sampling sequence.

[0050] Command energy characteristics are derived using the following formula:

[0051] (1)

[0052] In formula (1), This represents the instruction energy features extracted from the sampled sequence. Indicates the length of the sampling sequence. Represents the first in the sampling sequence Each sampling point. The control logic of formula (1) is based on the "average power calculation of discrete sampling sequence". Through the process of "sampling sequence amplitude squared → full sequence energy summation → length normalization", stable instruction energy features are extracted from the speech sampling sequence, providing a quantitative basis for subsequent semantic recognition and instruction validity screening.

[0053] The first in the sampling sequence sampling points This can be derived from the following formula:

[0054] (2)

[0055] In formula (2), This indicates the audio acquisition and conversion function of the audio acquisition device. Represents analog audio signals. Indicates the sampling frequency. The sampling index is indicated. The control logic of formula (2) is based on "discretization sampling of continuous analog audio signals + analog-to-digital conversion". Through the process of "continuous signal acquisition → time discretization → amplitude sampling → analog-to-digital conversion", the continuous speech commands in nature are converted into a digital sampling sequence that can be processed by a computer, providing a basic digital signal input for subsequent command energy feature extraction and semantic recognition.

[0056] The command energy feature extraction unit captures the user's speech signal using a built-in microphone as an audio acquisition device, converting it into a digital sampling sequence. This digital sampling sequence is typically time-series data, containing amplitude and frequency information of the speech. Extracting command energy features involves Fourier transform to obtain spectral features, or using Mel-frequency cepstral coefficients to capture the perceptual characteristics of the speech. Specifically, in the extraction stage, the sampling sequence is first preprocessed, such as removing noise and performing endpoint detection, to isolate valid speech segments. Then, a window function is applied for frame segmentation, and basic features such as short-time energy and zero-crossing rate are calculated for each frame. Further advanced features such as fundamental frequency and formants are extracted. These features are combined into a feature vector for subsequent intent recognition. In this way, by progressively building the feature set, a smooth transition from raw audio to analyzable data is ensured, improving the robustness of the system.

[0057] The keyword extraction unit is used to generate an intent vector by encoding the instruction energy features using a deep neural network model, and then extracts search keywords based on the intent vector.

[0058] The intent vector is derived using the following formula:

[0059] (3)

[0060] In formula (3), Represents the intent vector. Indicates the command energy characteristics, Represents the parameters of a deep neural network model. This represents a deep neural network. The control logic of formula (3) is based on "deep encoding of instruction energy features → generation of high-dimensional intent vectors". Through the process of "input of low-dimensional energy features → nonlinear mapping of deep neural network → output of high-dimensional intent vectors", the energy features of speech instructions are transformed into high-dimensional vectors that encode the core intent, providing a semantic quantitative basis for the accurate extraction of subsequent search keywords.

[0061] The match score for search keywords is calculated using the following formula:

[0062] (4)

[0063] In formula (4), Indicates search keywords Match score, This indicates the transpose of the intent vector. Indicates search keywords The corresponding embedding vector. The control logic of formula (4) is based on "quantification of the dot product similarity between intent vector and keyword embedding". Through the process of "vector transpose → dot product calculation → matching score output", the semantic matching degree between intent vector and search keyword embedding is quantified, providing a quantitative basis for accurately extracting search keywords.

[0064] The keyword extraction unit employs a deep neural network model to encode the energy features of the instruction, generating an intent vector. For example, a hybrid model combining convolutional neural networks and recurrent neural networks is used to extract local patterns through multi-layer convolution of the feature vector. Then, a long short-term memory unit captures temporal dependencies, ultimately outputting a fixed-dimensional intent vector. This intent vector represents the semantic essence of the user's instruction; for instance, "turn up the lights" corresponds to a vector indicating brightness adjustment. The process of extracting search keywords based on this intent vector involves focusing on key dimensions within the vector using an attention mechanism and mapping them to a predefined keyword library. For example, words like "brightness" and "adjust" are inferred from the vector, thus forming a search query. This seamless integration of encoding and extraction allows the system to extract concrete, actionable elements from abstract representations, avoiding the limitations of direct keyword matching and improving the accuracy of intent understanding in practical applications.

[0065] The semantic result output unit is used to locate the parameter index or generate optimization strategies based on the search keywords, and output semantic results related to network parameter query or signal optimization based on the parameter index or optimization strategies.

[0066] The semantic result is obtained using the following formula:

[0067] (5)

[0068] In formula (5), Indicates semantic result, Indicates the parameter index. Indicates the optimization strategy. The output function is used to generate network parameter queries or signal optimization results. The control logic of formula (5) is based on the semantic mapping of parameter index / optimization strategy. Through the process of "input type judgment → output function mapping → natural language result generation", the machine-understandable parameter index or optimization strategy is transformed into a natural language semantic result that the user can directly understand, thus completing the closed-loop output of voice interaction.

[0069] Parameter Index This can be derived from the following formula:

[0070] (6)

[0071] In formula (6), Indicates search keywords, Indicates the first A set of parameters Represents the similarity function. This indicates the position of the maximum value. The control logic of formula (6) is based on "similarity matching between search keywords and parameter sets + optimal index positioning". Through the process of "traversing candidate parameter sets → similarity measurement → maximum similarity index output", it quickly locates the pre-trained parameter set that best matches the current task, providing an accurate index basis for loading model parameters of 5G MIFI devices.

[0072] Optimization strategy This can be derived from the following formula:

[0073] (7)

[0074] In formula (7), Indicates weight, Represents the loss function. Indicates search keywords, Represents network parameters, This indicates an optimization operation. Indicates all Corresponding weights The control logic of formula (7) is based on "multi-task weighted loss aggregation + optimizer solving for optimal strategy". Through the process of "task weight allocation → multi-dimensional loss weighted summation → optimizer minimizing loss → outputting executable optimization strategy", it generates network parameters or signal optimization schemes that accurately match the search keywords, providing a quantitative decision basis for resource scheduling and performance optimization of 5G MIFI devices.

[0075] The semantic result output unit locates parameter indexes or generates optimization strategies based on search keywords. In telecommunications network optimization scenarios, if the keyword is "weak signal," the system can query parameter indexes in the database, such as the index values ​​of base station power levels or antenna tilt angles, and then generate optimization strategies such as increasing power output or adjusting the tilt angle to improve coverage. Specifically, this process involves matching keywords with a network parameter library, using a rule engine or machine learning model to evaluate the current network status, for example, by analyzing user location and signal strength logs to calculate optimization paths, such as prioritizing the adjustment of parameters of nearby base stations to minimize interference. Finally, based on these, the system outputs semantic results, such as natural language feedback like "signal strength has been increased to 85%." This logical progression from keywords to strategies not only automates parameter queries but also brings real-time response and efficient resource utilization to signal optimization, ensuring network stability and user satisfaction.

[0076] Furthermore, the 5G MiFi device integrating voice and visual recognition provided in this embodiment includes an optimization suggestion feedback data acquisition module 20 comprising an instruction type label determination unit, an encrypted request message encapsulation unit, a network performance bottleneck feature acquisition unit, and an optimization suggestion feedback data generation unit. The instruction type label determination unit is used to parse the intent slot information contained in the semantic results and compare the intent slot information with a preset instruction set to determine the instruction type label.

[0077] Intended slot information is obtained using the following formula:

[0078] (8)

[0079] In formula (8), Indicates the intended slot information. Represents an analytic function. The semantic result is represented. The control logic of formula (8) is based on "structured parsing of natural language semantic results → extraction of intent slots". Through the process of "semantic result input → structured extraction of parsing function → slot information output", the user-friendly natural language semantic results are transformed into structured intent slots that can be accurately processed by the machine, providing fine-grained semantic basis for the determination of subsequent instruction type labels.

[0080] Instruction type labels are derived using the following formula:

[0081] (9)

[0082] In formula (9), Indicates the instruction type label, Indicates the first Each comparison score The label corresponding to the maximum value is taken. The control logic of formula (9) is based on the core of "comparison score sorting between intent slot and preset instruction set → optimal instruction type label positioning". Through the process of "multi-label comparison score calculation → maximum value index positioning → instruction type label output", the type of the current instruction is accurately determined, providing a clear instruction classification basis for subsequent encryption request encapsulation and performance bottleneck analysis.

[0083] No. Each comparison score This can be derived from the following formula:

[0084] (10)

[0085] In formula (10), Represents the alignment function. This represents a preset instruction set. The control logic of formula (10) is based on "quantitative matching of structured intent slots and preset instruction templates → comparison score generation". Through the process of "slot and template alignment → comparison function quantification of matching degree → score output", it provides fine-grained matching basis for determining instruction type labels.

[0086] The instruction type label determines the intent slot information contained in the semantic results parsed by the unit. Intent slot information is a key element extracted from user instructions by the semantic understanding algorithm. For example, when a user queries "check current network speed," the intent slots include specific slot values ​​such as "check" and "network speed." These slots act like variables filling the intent framework, helping the system capture the core components of the user's intent. Specifically, this parsing involves decomposing the semantic results into structured data, such as using natural language processing tools to fill the slots in the results. Assuming the semantic result is a JSON object containing the intent name and related parameters, the slot information is extracted by iterating through these parameters. For example, if the intent name is "query parameters," the slot information might include "parameter type: speed," "location: current," etc., thus providing basic data for subsequent comparisons. The intent slot information is compared with a pre-defined instruction set to determine the instruction type label. This is achieved through a matching algorithm. The pre-defined instruction set is a predefined instruction library containing various categories such as query and optimization. Each category is associated with a set of keywords or patterns. During the comparison, the system calculates the similarity between the slot information and the instruction set. For example, cosine similarity is used to measure the degree of matching between the vectorized slot and the instruction template. If the similarity exceeds a threshold such as 0.8, the defined label "network parameter query category" is assigned. This process ensures the accurate classification of instructions and avoids mishandling of ambiguous instructions.

[0087] The encrypted request message encapsulation unit is used to read the underlying registers to generate the original network status data and encapsulate it into an encrypted request message if the instruction type label is a network parameter query category.

[0088] The following formula is used to define the condition for reading the underlying register:

[0089] (11)

[0090] In formula (11), Indicates a conditional trigger flag. Indicate the instruction type label, Indicates the category of network parameter query. The matching function is 1 when the instruction type label matches and 0 otherwise. The control logic of formula (11) is based on "precise matching of instruction type label and target category → generation of binary trigger flag". Through the process of "instruction type comparison → binarization of matching result → output of trigger flag", it provides a clear execution trigger signal for the encryption request message encapsulation unit, ensuring that the underlying register read operation is only executed when the instruction type is "network parameter query", thus ensuring system efficiency and security.

[0091] The raw network state data is encapsulated into an encrypted message using the following formula:

[0092] (12)

[0093] In formula (12), This indicates an encrypted request message. This indicates the encapsulation of encryption functions. This represents the original network state data. The control logic of formula (12) is based on "structured encapsulation of original network state data + encryption processing → secure message generation". Through the process of "original data input → protocol format encapsulation → encryption algorithm processing → encrypted message output", sensitive underlying network data is transformed into encrypted request messages that can be transmitted securely, thus ensuring the communication security between 5G MIFI devices and the backend server.

[0094] Raw network state data This can be derived from the following formula:

[0095] (13)

[0096] In formula (13), Represents the original network state data. Represents a generating function. This indicates a read operation. The control logic of formula (13) is based on "reading the raw data of the underlying hardware register → generating the structured network state". Through the process of "hardware trigger reading → obtaining binary raw data → parsing to generate structured data → outputting the raw state", the binary raw data of the underlying hardware is transformed into structured network state data that can be processed by the upper-layer software, providing accurate raw input for subsequent encrypted message encapsulation.

[0097] The encrypted request message encapsulation unit is used to read the underlying registers to generate the original network status data and encapsulate it into an encrypted request message if the instruction type label is a network parameter query category. Here, the underlying registers refer to the storage units at the device hardware level, such as the configuration registers in a router, which store real-time network metrics such as bandwidth utilization and latency values. When generating data, the system calls the API (Application Programming Interface) to read these values ​​from the registers. For example, if the bandwidth is 100Mbps and the latency is 50ms, the system serializes these data into binary format and encapsulates them into a message using the AES encryption algorithm. The message structure includes a header, a data body, and a checksum to ensure secure transmission.

[0098] The network performance bottleneck feature acquisition unit is used to send encrypted request messages to the cloud server, and the cloud server inputs the large model to obtain the network performance bottleneck features.

[0099] The characteristics of network performance bottlenecks are derived using the following formula:

[0100] (14)

[0101] In formula (14), Indicates network performance bottleneck characteristics. Indicates a large model input. The control logic of formula (14) is based on "multi-dimensional reasoning of the cloud-based large model → accurate identification of network performance bottlenecks". Through the process of "encrypted message upload → cloud decryption and parsing → multi-dimensional analysis of the large model → bottleneck feature output", it can extract hidden performance bottlenecks from the original network state data and provide accurate problem location basis for the generation of subsequent optimization suggestions.

[0102] Large model input This can be derived from the following formula:

[0103] (15)

[0104] In formula (15), This indicates that the message is received in the cloud. The input function is represented. The control logic of formula (15) is based on "parsing and standardizing the received messages in the cloud → generating inputs that can be processed by the large model". Through the process of "decrypting encrypted messages → extracting core data → standard preprocessing → model input and output", the heterogeneous messages received in the cloud are transformed into standardized inputs that can be directly processed by the large model, providing a high-quality data foundation for subsequent network performance bottleneck analysis.

[0105] Cloud message reception This can be derived from the following formula:

[0106] (16)

[0107] In formula (16), Indicates cloud server. The transmission function is represented. The control logic of formula (16) is based on "5G end-to-end transmission of encrypted request messages → reliable cloud reception". Through the process of "message encapsulation → 5G wireless transmission → cloud verification and reception → message delivery", the encrypted request messages generated by the 5G MIFI device are securely and with low latency transmitted to the cloud server, providing a reliable data channel for subsequent large model analysis.

[0108] The network performance bottleneck feature acquisition unit sends encrypted request messages to the cloud server. The cloud server inputs the large model to obtain network performance bottleneck features. The large model usually refers to a pre-trained language model such as GPT (Generative Pre-trained Transformer) series. When running on the cloud server, the decrypted message data is used as input prompts, such as "Analyze the following network conditions: bandwidth 100Mbps, latency 50ms, find the bottleneck". The large model processes through a multi-layer Transformer structure to infer bottleneck features such as "congestion caused by high latency". This feature extraction depends on the model's semantic understanding ability and combines massive training data to identify patterns.

[0109] The optimization suggestion feedback data generation unit is used to generate optimization suggestion feedback data containing configuration adjustment parameters based on network performance bottleneck characteristics.

[0110] The following formula is used to generate optimization suggestion feedback data containing optimal configuration adjustment parameters by minimizing the residuals:

[0111] (17)

[0112] In formula (17), This indicates the feedback data for optimization suggestions. This indicates that the system adjusts the mapping matrix. This represents a configuration adjustment parameter vector. The control logic of formula (17) is based on "least square optimization → optimal configuration adjustment parameter generation". Through the process of "bottleneck features and mapping matrix input → residual objective function construction → least square optimal solution solution → optimization suggestion feedback data output", it generates a configuration adjustment scheme that can accurately alleviate network performance bottlenecks and provides executable optimization instructions for 5G MIFI devices.

[0113] The optimization suggestion feedback data generation unit generates optimization suggestion feedback data containing configuration adjustment parameters based on the characteristics of network performance bottlenecks. This generation process involves a rule-based or learning-based strategy module. For example, if the bottleneck is "high latency", the parameter library is queried to generate adjustment parameters such as "increase the buffer size to 2MB", and then it is encapsulated into feedback data such as a suggestion list in JSON format, including parameter name, adjustment value and expected effect. This ensures that the feedback data can be directly used for device configuration updates. In this way, the logical chain from bottleneck analysis to suggestion generation remains coherent, supporting dynamic network optimization.

[0114] Furthermore, the 5G MiFi device integrating voice and visual recognition provided in this embodiment includes a preliminary evaluation result acquisition module 30 comprising a spatiotemporal synchronization data packet generation unit, a user gesture command encoding acquisition unit, a legitimate control request signal generation unit, and a preliminary evaluation result output unit. The spatiotemporal synchronization data packet generation unit is used to receive optimization suggestion feedback data sent from the cloud and visual acquisition data collected by the visual module, and to align the optimization suggestion feedback data and visual acquisition data to generate a spatiotemporal synchronization data packet.

[0115] The following formula is used to align the optimization suggestion feedback data with the visual acquisition data to generate a synchronized data packet:

[0116] (18)

[0117] In formula (18), This represents a time-space synchronization data packet. This indicates the feedback data for optimization suggestions. This represents visually acquired data. The alignment function is represented. The control logic of formula (18) is based on "precise alignment of spatiotemporal dimensions of multimodal data → generation of synchronous data packets". Through the process of "timestamp matching → spatial coordinate association → structured encapsulation", the optimization suggestion feedback data sent from the cloud and the local visual acquisition data are bound to a unified spatiotemporal coordinate system to generate synchronous data packets containing time, space, instructions and visual information, providing accurate input for multimodal fusion for subsequent preliminary evaluation.

[0118] The spatiotemporal synchronization data packet generation unit receives optimization suggestion feedback data from the cloud and visual acquisition data collected by the vision module. The feedback data consists of configuration adjustment parameters generated by the cloud-based large model based on previously uploaded network conditions, such as specific values ​​for bandwidth allocation strategies or latency optimization schemes. The visual acquisition data consists of user image sequences captured in real time by the vision module, such as a camera. This data includes video frames of user gestures in the smart home environment. Specifically, the alignment process involves a timestamp synchronization mechanism, matching the generation time of the optimization suggestion feedback data with the acquisition time of the visual acquisition data. For example, assuming the timestamp of the optimization suggestion feedback data is T1 and the visual data is a continuous frame sequence from T0 to T2, a linear interpolation algorithm is used to adjust the time axis so that the two are aligned within a common time window, thereby generating a spatiotemporal synchronization data packet. This packet structure includes a header identifier, a synchronization timestamp, a feedback parameter block, and an image frame block, ensuring data consistency during subsequent processing. For example, in a smart conferencing system, the cloud feedback data might be "adjust the video stream bitrate to 2Mbps," and the visual data might be a video of the user waving their hand. The data packet generated after alignment can support real-time interactive control, avoiding misinterpretation of commands due to timing discrepancies.

[0119] The user gesture command encoding acquisition unit is used to extract image frames from the spatiotemporal synchronization data packet and calculate the user gesture command encoding.

[0120] User gesture command encoding is derived using the following formula:

[0121] (19)

[0122] In formula (19), Indicates the encoding of user gesture commands. Gestures Match score, The formula (19) represents the gesture category. It means that the final user gesture instruction code is obtained by converting the gesture category with the highest score into a one-hot code. The control logic of the formula (19) is based on "gesque image category matching → one-hot code instruction generation". Through the process of "image feature extraction → multi-category matching score calculation → highest score category positioning → one-hot code conversion", the user's visual gesture is converted into an instruction code that the machine can directly process, providing accurate user intent basis for the generation of subsequent legitimate control requests.

[0123] gesture Matching score This can be derived from the following formula:

[0124] (20)

[0125] In formula (20), Indicates the length of the image frame sequence. Indicates the first Image frame, Gestures The template features, this formula means that the gesture matching score is calculated by averaging the cosine similarity between the frame sequence and the template. The control logic of formula (20) is based on "averaging the cosine similarity between the continuous image frame sequence and the gesture template → generating robust matching scores". Through the process of "multi-frame feature extraction → single-frame similarity calculation → sequence average score → matching score output", a stable and accurate gesture matching score is generated, which provides a reliable similarity basis for subsequent gesture command encoding.

[0126] No. Image Frame This can be derived from the following formula:

[0127] (twenty one)

[0128] In formula (21), Indicates the size of spatial dimensions. Indicates the spatial location of the spatiotemporal synchronization data packet. and time The value at that location, Indicates the first The timestamp, this formula means that a single frame image is extracted by the spatial averaging of spatiotemporal data packets. The control logic of formula (21) is based on "spatial dimension averaging of spatiotemporal synchronization data packets → single frame visual image extraction". Through the process of "timestamp locking → spatial data aggregation → average noise reduction → single frame output", a clear and stable single frame visual image is extracted from the data packets containing spatiotemporal information, providing high-quality visual input for subsequent gesture feature matching.

[0129] The user gesture command encoding acquisition unit extracts image frames from spatiotemporal synchronization data packets and calculates the user gesture command encoding. The process of extracting image frames first parses the image blocks of the data packets, using tools such as the OpenCV library to read them frame by frame, for example, extracting a sequence of 30 frames per second from the packet. Then, these image frames are preprocessed, such as grayscale conversion and edge detection, to highlight the gesture outline. Specifically, calculating the gesture command encoding involves inference of a convolutional neural network model. For example, the preprocessed frames are input into a pre-trained CNN (Convolutional Neural Network) model. This CNN model extracts features such as finger position and trajectory vectors through multiple convolutions, and then outputs encoded values ​​through fully connected layers. For example, "waving upwards" corresponds to encoding 01, and "clenching a fist" corresponds to 02. This encoding is based on a standardized representation of the gesture library to ensure accurate capture of user intent.

[0130] The legitimate control request signal generation unit is used to generate a legitimate control request signal if the user's gesture command encoding matches the preset permission control requirements.

[0131] A legitimate control request signal is derived using the following formula:

[0132] (twenty two)

[0133] In formula (22), This indicates a legitimate control request signal. This formula represents the pre-defined access control requirements and describes the framework for generating signals under precise matching conditions, when user gesture commands are encoded. With pre-defined access control requirements When they are equal, the output is 1 to indicate the generation of a valid signal; otherwise, it is 0. The control logic of formula (22) is based on "precise matching of gesture command encoding and permission requirements → generation of valid control requests". Through the process of "permission rule matching → binary signal output → execution process triggering", the security verification of control requests is realized to ensure that only gesture commands that meet the preset permissions can trigger the device optimization operation and ensure the operational security of 5G MIFI devices.

[0134] The legitimate control request signal generation unit generates a legitimate control request signal if the user's gesture command encoding matches the preset permission control requirements. These permission control requirements are a pre-defined set of rules, such as user authentication and gesture whitelists, for example, encoding 01 only allows administrator-level users to execute. Specifically, the matching process is implemented through a comparison algorithm, comparing the extracted encoding with each entry in the rule set. If a match is successful (e.g., the encoding matches and the user ID is in the authorized list), a legitimate control request signal is generated. This legitimate control request signal is a binary flag carrying the request type, such as "network configuration update," for subsequent module responses.

[0135] The preliminary evaluation result output unit is used to construct a scene topology diagram based on legitimate control request signals, and to map the optimization suggestion feedback data to the scene topology diagram to output the preliminary evaluation results of multi-user network access permissions.

[0136] The preliminary assessment results of multi-user network access permissions are obtained using the following formula:

[0137] (twenty three)

[0138] In formula (23), This indicates the preliminary assessment results of multi-user network access permissions. Indicates the total number of evaluation dimensions. Indicates the first Wei Zai Mapping data at location, The formula (23) represents the sigmoid activation function and is used to output the preliminary evaluation results of multi-user network access permissions based on the mapping data. The control logic of the formula (23) is based on "normalization and averaging of multi-dimensional mapping data → generation of comprehensive permission evaluation score". Through the process of "acquiring multi-dimensional mapping data → sigmoid normalization → arithmetic mean aggregation → output of evaluation results", a comprehensive score between 0 and 1 is generated, providing a quantitative preliminary evaluation basis for the compliance of multi-user network access permissions.

[0139] No. Wei Zai Mapping data at location This can be derived from the following formula:

[0140] (twenty four)

[0141] In formula (24), This indicates that the optimization suggestions feedback data is in The value at the location, The scene topology diagram is shown in Edge weight at position, The formula represents element-wise multiplication and is used to map the optimization suggestion feedback data to the scene topology diagram. The control logic of formula (24) is based on "element-wise weighted combination of optimization suggestions and scene topology → generation of scene-aware mapping data". Through the process of "obtaining the original value of optimization suggestions → matching the edge weight of scene topology → element-wise weighted multiplication → outputting mapping data", the optimization suggestion feedback data is bound to the positional importance of the scene topology, generating quantitative mapping data containing scene awareness, and providing accurate scene-based input for subsequent multi-person network access permission assessment.

[0142] Scene topology diagram in Edge weight at position This can be derived from the following formula:

[0143] (25)

[0144] In formula (25), Indicates the first The first legitimate control request signal and the first The squared Euclidean distance of a valid control request signal Indicates the first A legitimate control request signal, Indicates the first A legitimate control request signal, The scale parameter of the Gaussian kernel is represented by the formula used to construct the scene topology graph based on the legitimate control request signal. The control logic of formula (25) is based on "similarity measurement of legitimate control request signal → Gaussian kernel weighted edge weight generation". Through the process of "signal difference calculation → Gaussian kernel mapping → weight output", the adjacency relationship of the scene topology is constructed based on the legitimate control request of the user, reflecting the similarity and connection tightness of user operations in multi-user networking.

[0145] The preliminary evaluation result output unit constructs a scene topology graph based on legitimate control request signals and maps optimization suggestion feedback data to the scene topology graph to output a preliminary evaluation result of multi-user network access permissions. The scene topology graph is a graphical representation, for example, using a graph theory model where nodes represent devices such as routers and terminals, and edges represent connections. It is dynamically constructed after being triggered by signals, such as adding weights between nodes based on the current network load. Specifically, the mapping process applies feedback data, such as "bitrate adjustment," to the edge attributes of the scene topology graph. For example, it updates the edge values ​​from the cloud to the terminal with optimized parameters. Then, it calculates connectivity and permission distribution using graph traversal algorithms such as BFS, outputting an evaluation result such as "User A can access shared resources with a high permission level." This result supports preliminary decision-making in multi-user collaborative scenarios, ensuring the secure implementation of network optimization.

[0146] Furthermore, the 5G MiFi device integrating voice and visual recognition provided in this embodiment includes a device connection list acquisition module 40 comprising a device access demand load calculation unit, an available spectrum resource block filtering unit, a multi-link concurrent access configuration table generation unit, and a device connection list acquisition unit. The device access demand load calculation unit is used to parse the preliminary evaluation results of multi-user network access permissions to calculate the device access demand load, which is determined based on the number of terminal requests and the service bandwidth requirement value.

[0147] The load requirement for device access is calculated using the following formula:

[0148] (26)

[0149] In formula (26), Indicates the load demand for device access. This indicates the preliminary assessment results of multi-user network access permissions. Indicates the number of terminal requests. The formula (26) indicates that the service bandwidth demand value is determined by multiplying the initial assessment result, the number of terminal requests, and the service bandwidth demand value. The control logic of formula (26) is based on "multi-factor coupling quantification → accurate assessment of device access demand load". Through the process of "compliance weighting → multiplication of scale and bandwidth → load value output", it combines the permission compliance, access scale, and bandwidth demand of multiple network users to generate a load value that reflects the effective demand intensity of network resources, providing a quantitative basis for subsequent spectrum resource screening and access configuration.

[0150] When the device access demand load calculation unit parses the preliminary assessment results of multi-user network access permissions, it first needs to calculate the device access demand load. This load is essentially a quantitative indicator of the system's consumption of network resources, determined by integrating the number of terminal requests and the service bandwidth requirement. Specifically, the number of terminal requests refers to the total number of connection requests from different devices within a specific time window, such as per minute. For example, in a smart office environment, multiple employee terminals may simultaneously initiate video conferencing requests, potentially reaching 50. The service bandwidth requirement is calculated based on the type of each request; for instance, high-definition video requires 2Mbps of bandwidth, while low-definition data transmission only requires 0.5Mbps. By summing these values, the total load is obtained. For example, if there are 30 high-definition requests and 20 low-definition requests, the total load is 30*2 + 20*0.5 = 70Mbps. This calculation process ensures the accuracy of the assessment results and supports subsequent resource allocation decisions.

[0151] The available spectrum resource block filtering unit is used to activate the wireless connection extension mechanism to filter available spectrum resource blocks if the device access demand load is greater than the single node access capacity threshold. Available spectrum resource blocks are extracted from the spare radio frequency channel.

[0152] The activation criteria for the wireless connectivity extension mechanism are defined by the following formula:

[0153] (27)

[0154] In formula (27), This indicates that the wireless connectivity extension mechanism is activated. This represents the single-node access capacity threshold. The formula describes the load demand when devices are connected. Exceeding the single-node access capacity threshold The extension mechanism is activated when the time is right, otherwise it is not activated. The control logic of formula (27) is based on “load and capacity threshold comparison → wireless connection extension mechanism triggering”. Through the process of “load and capacity acquisition → threshold judgment → activation flag output”, it dynamically determines whether to enable the available spectrum resources of the backup radio frequency channel to ensure the elastic expansion of network capacity in multi-user networking scenarios.

[0155] The following formula is used to describe the extraction of available spectrum resource blocks from spare radio frequency channels:

[0156] (28)

[0157] In formula (28), This indicates the extracted available spectrum resource blocks. Indicates a spare radio frequency channel. Indicates a set of backup radio frequency channels. Indicates the spare radio frequency channel The selected indicator function. The control logic of formula (28) is based on "availability judgment of backup radio frequency channel → quantitative extraction of available spectrum resource block". Through the process of "backup channel traversal → availability judgment → indicator function binarization → summation and statistical resource quantity", it accurately filters and quantifies the available spectrum resources in the backup radio frequency channel, providing clear resource support for multi-link concurrent access configuration.

[0158] The available spectrum resource block filtering unit is used to filter available spectrum resource blocks if the calculated device access demand load exceeds the single-node access capacity threshold (e.g., the threshold is set at 50Mbps while the actual load is 70Mbps). This activates the wireless connection extension mechanism to filter available spectrum resource blocks. This wireless connection extension mechanism is a dynamic resource management strategy that extracts unused frequency bands from spare radio frequency channels. Specifically, spare radio frequency channels include the 2.4GHz and 5GHz frequency bands reserved by the system. The filtering process involves scanning algorithms to detect signal strength and interference levels, for example, prioritizing channel blocks with interference below -80dBm. Each block is 20MHz in size. This extraction ensures that signal conflicts are avoided when extending connections, improving network stability.

[0159] The multi-link concurrent access configuration table generation unit is used to generate a multi-link concurrent access configuration table based on available spectrum resource blocks. The multi-link concurrent access configuration table is used to perform device registration and link migration.

[0160] The multi-link concurrent access configuration table is derived using the following formula:

[0161] (29)

[0162] In formula (29), This represents the optimal multi-link concurrent access configuration table. This indicates the total number of available spectrum resource blocks. This represents the utility function for matching resource blocks with configuration tables. Indicates the first One available spectrum resource block, The formula (29) represents the candidate configuration table. This formula describes the optimization process of generating a multi-link concurrent access configuration table based on available spectrum resource blocks. The control logic of formula (29) is based on "maximizing the global utility of multi-spectrum resources → generating the optimal multi-link configuration table". Through the process of "candidate configuration traversal → resource-configuration utility calculation → comprehensive score averaging → optimal scheme selection", a multi-link concurrent access configuration table that maximizes the overall network utility is generated from the available spectrum resource blocks, providing a precise resource allocation scheme for device registration and link migration.

[0163] The multi-link concurrent access configuration table generation unit generates a multi-link concurrent access configuration table based on the selected available spectrum resource blocks. This multi-link concurrent access configuration table is a structured data structure used to guide device registration and link migration operations. Specifically, the multi-link concurrent access configuration table includes columns such as resource block ID, allocated device ID, and migration priority. For example, block A is allocated to terminal 1, with higher priority for faster registration. When a new device registers, the table records its MAC address and authentication key. Link migration switches the device from a congested channel to an idle block when the load is too high. This process is achieved through table updates to ensure seamless connectivity and avoid service interruption.

[0164] The device connection list acquisition unit is used to update the mapping relationship based on the multi-link concurrent access configuration table to obtain a dynamically adjusted device connection list.

[0165] The dynamically adjusted device connection list is derived using the following formula:

[0166] (30)

[0167] In formula (30), This indicates the dynamically adjusted list of connected devices. Represents the original mapping relationship. Represents the mapping transformation function. The control logic of formula (30) is based on "dynamic update of mapping relationship driven by optimal configuration table → generation of real-time device connection list". Through the process of "original mapping formatting → optimal configuration guidance adjustment → dynamic list output", the decision of multi-link concurrent access is transformed into an executable device-link binding list, providing a direct execution basis for device registration and link migration.

[0168] The device connection list acquisition unit updates the mapping relationship based on the multi-link concurrent access configuration table to obtain a dynamically adjusted device connection list. This update involves reconstructing the correspondence between devices and resource blocks. Specifically, the mapping relationship is initially based on static allocation, but is adjusted through table data such as migration records, for example, moving terminal 2 from block A to block B. The dynamically adjusted device connection list then lists the currently connected channels and statuses of all devices. This dynamically adjusted device connection list supports real-time monitoring, ensuring access optimization after permission assessment in multi-user network scenarios, achieving efficient resource utilization and technical effects such as reduced latency and improved user experience.

[0169] Furthermore, the 5G MiFi device integrating voice and visual recognition provided in this embodiment includes a complex scene recognition conclusion acquisition module 50 comprising a voiceprint spectrum feature conversion unit, a dual-modal feature acquisition unit, a scene semantic label determination unit, and a complex scene recognition conclusion acquisition unit. The voiceprint spectrum feature conversion unit is used to acquire the real-time signal strength and voice command content in the device connection list and convert the voice command content into voiceprint spectrum features.

[0170] Real-time signal strength is obtained using the following formula:

[0171] (31)

[0172] In formula (31), This represents the real-time signal strength vector in the device connection list, and is a... A structured vector of dimensions, containing the current access devices at time [time]. The signal strength value is used to reflect network coverage quality and device connection stability. Indicates the first The real-time signal strength of each device is expressed in dBm (a larger value, such as −50dBm, indicates a stronger signal; a smaller value, such as −100dBm, indicates a weaker signal). This indicates the number of devices in the device connection list, representing the total number of terminals currently connected to the 5G MiFi device (e.g., M=2 indicates that there are 2 terminals connected). The timestamp indicates the acquisition time, ensuring that the signal strength is synchronized with the time of voice commands and visual data, and providing spatiotemporal consistency guarantee for multimodal feature fusion. The control logic of formula (31) is based on "vector encapsulation of real-time signal strength of multiple devices → generation of network environment perception data". Through the process of "device list traversal → single device signal acquisition → vector structured encapsulation → real-time perception output", the real-time signal strength of all access devices is collected in a unified manner, providing accurate environmental perception basis for subsequent dual-modal feature fusion and complex scene recognition.

[0173] The following formula is used to convert the speech command content into spectral power features through Fourier transform:

[0174] (32)

[0175] In formula (32), The characteristics of the acoustic spectrum are represented by angular frequencies. The function represents the energy distribution of the speech signal at different frequencies, reflecting the speaker's unique voiceprint characteristics such as pitch and timbre (e.g., the frequency peak position of the power spectrum is different when different people say "confirm"). The time-domain signal representing the content of a voice command is the sound wave pressure value that changes over time (such as the sound wave waveform of the "confirm" command, in Pa). It represents angular frequency. The complex exponential basis functions represent the Fourier transform, used to decompose a time-domain signal into different angular frequencies. The sine / cosine components. The continuous-time Fourier transform maps a time-domain signal to the frequency domain, yielding a complex spectrum containing amplitude and phase information. . This means taking the square of the modulus, which converts the complex spectrum into a power spectrum (the square of the amplitude), eliminating phase information and retaining only the correspondence between frequency and energy, making the characteristics more stable.

[0176] The control logic of formula (32) is based on "frequency domain decomposition of time-domain speech signal → generation of voiceprint power spectrum features". Through the process of "time domain signal acquisition → continuous Fourier transform frequency domain decomposition → modulus square power spectrum conversion → voiceprint feature output", the time domain sound wave of the speech command is transformed into the frequency domain power distribution reflecting the speaker's unique voiceprint, providing a stable voiceprint representation for dual-modal feature fusion and complex scene recognition.

[0177] The process of acquiring real-time signal strength and voice command content by the voiceprint spectrum feature conversion unit first involves capturing environmental data using built-in sensors. For example, in a smart home system, the device monitors the RSSI (Received Signal Strength Indicator) value of the Wi-Fi signal in real time via a wireless module. This RSSI value represents the received signal strength, usually quantized in dBm. Simultaneously, a microphone array records the user's voice commands, such as "turn on the lights." This acquisition ensures data synchronization, providing a foundation for subsequent processing. Specifically, when converting the voice command content into voiceprint spectrum features, a Fourier transform algorithm is needed to convert the time-domain speech signal into a frequency-domain representation. For example, Mel-frequency cepstral coefficients are extracted as feature vectors. In an industrial monitoring scenario, if the command is "check machine status," the system analyzes the fundamental frequency and formants of the speech to form a multi-dimensional spectrum. This conversion process includes pre-emphasis and window function application to highlight high-frequency components and reduce edge effects, thereby obtaining a more accurate voiceprint representation.

[0178] The dual-modal feature acquisition unit is used to generate enhanced texture features by combining visual diagnostic data if the real-time signal strength is lower than the threshold, and then aggregate the enhanced texture features with the acoustic spectrum features to obtain dual-modal features.

[0179] The following formula is used to define the triggering conditions for enhanced texture features:

[0180] (33)

[0181] In formula (33), This indicates a trigger indicator, which is a binary signal of 0 or 1. The threshold (e.g., -70dBm) is the critical value for determining whether the signal needs to be enhanced. The control logic of formula (33) is based on "threshold triggering of signal quality → adaptive visual feature enhancement". Through the process of "single device signal strength comparison → binary trigger indicator output → enhancement process start / skip", it automatically enhances visual texture features when the signal quality is poor, makes up for the shortcomings of speech recognition, and improves the robustness of dual-modal features.

[0182] The bimodal characteristics are derived using the following formula:

[0183] (34)

[0184] In formula (34), Representing bimodal features, it is a fused unified feature vector / tensor containing spatial information of visual texture and frequency domain information of voiceprint spectrum, used for semantic label determination in subsequent complex scenes. This represents an aggregation function, responsible for fusing features from two modalities. The enhanced texture features (from the vision module) are high-resolution visual features (such as clear gesture textures and scene edge information) that are generated after the enhancement is triggered. The voiceprint spectrum feature (derived from Fourier transform) is the frequency domain power spectrum feature of the speech command (reflecting the speaker's identity and speech content). The control logic of formula (34) is based on "multi-dimensional aggregation of visual and speech features → unified bimodal feature generation". Through the process of "enhanced texture and voiceprint feature acquisition → aggregation function fusion → unified feature output", it combines the enhanced texture information of the visual domain with the voiceprint spectrum information of the speech domain to generate a unified feature containing multimodal semantics, providing richer and more robust input for complex scene recognition.

[0185] Enhanced texture features are derived using the following formula:

[0186] (35)

[0187] In formula (35), Enhanced texture features are processed high-resolution, low-noise visual features (such as clear gesture edges and scene texture details) that contain rich spatial semantic information. This represents the generator function, which is a wrapper around the visual enhancement algorithm. Visual diagnostic data refers to raw visual acquisition data (such as gesture images and scene frames) and signal quality diagnostic information (such as trigger indicators). =1, RSSI value) combined, including "visual content + quality label", making the enhancement more targeted. The control logic of formula (35) is based on "adaptive enhancement driven by visual diagnostic data → high robust texture feature generation". Through the process of "diagnostic data input → generation function enhancement processing → enhanced texture output", when the signal quality is poor, the original visual data is subjected to targeted enhancement processing to generate clear, low noise texture features, providing reliable visual support for dual-modal fusion.

[0188] The dual-modal feature acquisition unit is used to generate enhanced texture features by combining visual diagnostic data if the real-time signal strength is below a threshold, such as a threshold set at -70dBm but the actual measurement is -80dBm. This threshold triggers the auxiliary mode. The visual data comes from images captured by a camera; for example, in autonomous vehicles, the diagnostic data is road surface images. Generating enhanced texture features involves using Gabor filters to extract edges and texture details, and then increasing contrast and sharpening to form an enhanced feature matrix. This process helps compensate for information loss when the signal is weak. When aggregating the enhanced texture features with the acoustic spectrum features to obtain dual-modal features, a feature fusion strategy is used. For example, the grayscale gradient of the visual texture and the acoustic spectrum are combined through cascading or attention mechanisms to form a unified vector representation. This aggregation includes dimension alignment and weighted summation to ensure the complementarity of the two modalities.

[0189] The scene semantic label determination unit is used to input bimodal features into the fault classification model to obtain the fault probability distribution, and determine the scene semantic label based on the bimodal features.

[0190] The following formula is used to input bimodal features into a fault classification model to obtain the fault probability distribution:

[0191] (36)

[0192] In formula (36), Representing the failure probability distribution, it is a dimensional vector ( (Number of fault / scenario categories). The weight matrix represents the fault classification model, which consists of parameters learned by the model during training. It is used to map bimodal features to the fault category space. This represents the bias vector, which is a parameter used during model training to adjust the baseline value of the probability for each class (to avoid probability bias caused by feature distribution shifts). The normalization exponential function represents the transformation of the unnormalized score (logits) after linear transformation into a probability distribution between 0 and 1, ensuring that the sum of the probabilities of all categories is 1. The control logic of formula (36) is based on "linear mapping of bimodal features → probability distribution normalization → scene fault probability output". Through the process of "feature linear transformation → Softmax normalization → probability distribution generation", the fused bimodal features are transformed into probability distributions of various faults / scenes, providing a quantitative confidence basis for the determination of scene semantic labels.

[0193] Scene semantic tags are derived using the following formula:

[0194] (37)

[0195] In formula (37), These represent scene semantic tags, which are indexes of scene categories and directly correspond to the semantic description of the scene, eliminating the need for manual interpretation of probabilities. The scene semantic classification weight matrix represents the parameters obtained during model training, used to map bimodal features to the scene category score space. This represents the bias vector, used to adjust the score baseline for each scene category, improving classification stability (avoiding score deviation caused by feature distribution shift). The category index is represented. The control logic of formula (37) is based on "linear classification of bimodal features → extraction of maximum score index → ​​generation of scene semantic label". Through the process of "linear mapping of features → calculation of category score → output of maximum score index", the fused bimodal features are directly mapped to a unique scene semantic label, providing clear semantic conclusions for complex scene recognition.

[0196] The scene semantic label determination unit inputs bimodal features into a fault classification model to obtain the fault probability distribution. It then uses a pre-trained neural network, such as a CNN-RNN hybrid model, for classification. For example, in a smart factory, after inputting features, the fault classification model outputs probability vectors for various faults, such as "equipment overheating." This distribution is normalized using a softmax function and inferred based on patterns learned from the training dataset. Specifically, determining scene semantic labels based on bimodal features involves semantic segmentation algorithms to extract labels. For example, if the label is "noise interference," this determination is achieved through clustering or embedding space analysis, connecting with the aforementioned probability distribution to provide contextual understanding.

[0197] The complex scene recognition conclusion acquisition unit is used to aggregate the fault probability distribution and scene semantic labels to generate a diagnostic decision vector, parse the diagnostic decision vector to determine the fault diagnosis output, and obtain the complex scene recognition conclusion.

[0198] The following formula is used to generate a diagnostic decision vector by weighted aggregation of the fault probability distribution and scene semantic labels:

[0199] (38)

[0200] In formula (38), The diagnostic decision vector is the final fusion decision output, which includes the weighted contribution of fault probability and the weighted contribution of scene semantics. Each element in the vector corresponds to the comprehensive score of different faults / scenes, which is directly used to trigger subsequent diagnostic actions (such as activating backup spectrum or adjusting transmission power). The fault probability weight matrix is ​​a parameter (usually a diagonal matrix) used for model training or business configuration to assign differentiated importance to different fault types (e.g., "weak signal" faults have a higher weight in multi-user network scenarios, while "normal" faults have a lower weight). This represents the scene semantic weight matrix, used to assign semantic labels to different scenes. Assign priority weights (e.g., the "outdoor occlusion" scenario has a higher weight because it needs to trigger backup spectrum expansion more). The control logic of formula (38) is based on "weighted fusion of fault probability and scenario semantics → generation of multi-dimensional diagnostic decision vector". Through the process of "double input weighted mapping → vector aggregation → decision vector output", it combines the quantitative confidence of fault probability with the priority information of scenario semantics to generate a vector containing multi-dimensional decision basis, providing clear quantitative support for fault diagnosis and action triggering in complex scenarios.

[0201] The conclusion for complex scene recognition is derived using the following formula:

[0202] (39)

[0203] In formula (39), The conclusion of complex scene recognition is the final judgment result of binarization (1 means "complex scene, optimization action needs to be triggered"; 0 means "normal scene, no intervention required"), which directly guides the device to perform preset actions (such as activating the backup spectrum and adjusting the transmission power). This represents the fault diagnosis output (from the diagnostic decision vector). The highest score, such as =0.82 (corresponding to the comprehensive score of "weak signal scenario + fault"), representing the risk intensity of the current scenario. The decision threshold represents a pre-configured critical value (e.g., ...). =0.7), used to distinguish between scenarios that "require intervention" and "do not require intervention". The indicator function returns 1 when the condition in parentheses is true and 0 when it is false. It is the key operator for converting continuous scores into discrete conclusions. The control logic of formula (39) is based on "threshold judgment of fault diagnosis score → binarized scene conclusion output". Through the process of "comparison of diagnosis score with threshold → binarization of indicator function → generation of scene conclusion", the continuous diagnosis score is converted into a clear scene recognition conclusion, which directly triggers the subsequent network optimization action.

[0204] Fault diagnosis output This can be derived from the following formula:

[0205] (40)

[0206] In formula (40), The first vector representing the diagnostic decision vector Quantity, is the first The comprehensive weighted score of each fault / scenario (e.g.) =0.82 represents the overall score for "weak signal fault". This represents the fault category index, corresponding to different fault / scenario types (e.g., =0 is normal. =1The signal is weak, =2 (high interference). This formula is used to analyze the diagnostic decision vector and determine the fault diagnosis output by taking the maximum value. The control logic of formula (40) is based on "extraction of the maximum score index of the diagnostic decision vector → location of the core fault type". Through the process of "traversal of decision vector → output of the maximum score index → ​​generation of fault diagnosis result", it accurately locates the fault type with the highest score from the multi-dimensional diagnostic decision vector, and provides a clear core fault basis for the identification conclusion of complex scenarios.

[0207] When the complex scene recognition conclusion acquisition unit aggregates the fault probability distribution and scene semantic labels to generate a diagnostic decision vector, it uses a vector concatenation method to fuse the two. For example, it merges the probability vector and the label encoding vector into a high-dimensional decision vector, which supports comprehensive evaluation. Specifically, the process of parsing the diagnostic decision vector to determine the fault diagnosis output and obtain the complex scene recognition conclusion involves parsing vector elements through threshold comparison or decision tree analysis. For example, the output might be "connection fault caused by signal interference." This conclusion integrates all information to achieve accurate recognition.

[0208] The 5G MiFi device integrating voice and visual recognition provided in this embodiment, compared with existing technologies, obtains user intent through voice semantic understanding, combines the current downlink rate and carrier aggregation status with a cloud-based large model to generate optimization suggestions, and integrates user gesture recognition results collected by the visual module for access control, ensuring that only authorized users can trigger multi-user network resource adjustments. When the initial assessment indicates that multiple users accessing the network exceeds a threshold, the wireless connection expansion mechanism is automatically activated and resources are dynamically allocated, ultimately forming an optimized device connection list. Furthermore, through dual-modal fusion of voice commands and visual diagnostic data, accurate fault diagnosis and identification are achieved in complex scenarios. This embodiment effectively solves the three core problems caused by concurrent access by multiple users: network congestion, uncontrolled access, and difficulty in fault location, significantly improving network security, stability, and user experience.

[0209] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A 5G MiFi device integrating voice and visual recognition, characterized in that, include: The semantic result acquisition module (10) is used to acquire audio signals from user voice input, acquire the audio signals through an audio acquisition device, and use a semantic understanding algorithm to parse the voice command content of the audio signals to obtain semantic results related to network parameter query or signal optimization. The optimization suggestion feedback data acquisition module (20) is used to determine the instruction type based on the semantic result. If the semantic result includes network parameter query, the current downlink rate and carrier aggregation status are extracted from the network communication module and uploaded to the cloud big model for processing through the end-cloud collaborative communication channel to obtain optimization suggestion feedback data. The preliminary evaluation result acquisition module (30) is used to analyze user gestures using a gesture recognition algorithm based on the visual acquisition data collected by the visual module in response to the feedback data of the optimization suggestion. If the user gestures match the permission control requirements, the feedback data is combined with the image scene perception to obtain the preliminary evaluation result of the multi-person network access permission. The device connection list acquisition module (40) is used to determine the device access requirements based on the preliminary evaluation results. If the preliminary evaluation results show that multiple users access the device more than the threshold, the wireless connection extension mechanism is activated, access resources are allocated, and a dynamically adjusted device connection list is obtained. The complex scene recognition conclusion acquisition module (50) is used to obtain real-time signal strength indicators from the device connection list, integrate voice command content and visual diagnostic data using a dual-modal fusion algorithm, determine fault diagnosis output, and obtain complex scene recognition conclusions. The complex scene recognition conclusion acquisition module (50) includes: The voiceprint spectrum feature conversion unit is used to obtain the real-time signal strength and voice command content in the device connection list, and convert the voice command content into voiceprint spectrum features. Real-time signal strength is obtained using the following formula: ; in, This represents the real-time signal strength vector in the device connection list. Indicates the first Real-time signal strength of each device This indicates the number of devices in the device connection list. Indicates the collection timestamp; The following formula is used to convert the speech command content into spectral power features through Fourier transform: ; in, Represents the spectral characteristics of the voiceprint. The time-domain signal representing the content of the voice command. Represents angular frequency. Let represent the complex exponential basis functions of the Fourier transform. Represents the continuous-time Fourier transform. This indicates taking the square modulo 1. A dual-modal feature acquisition unit is used to generate enhanced texture features by combining visual diagnostic data if the real-time signal strength is lower than a threshold, and to aggregate the enhanced texture features with the acoustic spectrum features to obtain dual-modal features. The following formula is used to define the triggering conditions for enhanced texture features: ; in, Indicates trigger indicator, Indicates the threshold; The bimodal characteristics are derived using the following formula: ; in, Indicates bimodal characteristics, Represents aggregate functions, Indicates enhanced texture features, Indicates the spectral characteristics of the voiceprint; Enhanced texture features are derived using the following formula: ; in, Represents a generating function. Represents visual diagnostic data; The scene semantic label determination unit is used to input the bimodal features into the fault classification model to obtain the fault probability distribution, and determine the scene semantic label based on the bimodal features; The following formula is used to input bimodal features into a fault classification model to obtain the fault probability distribution: ; in, Represents the failure probability distribution. This represents the weight matrix of the fault classification model. This represents the bias vector. This represents the normalized exponential function; Scene semantic tags are derived using the following formula: ; in, Represents scene semantic tags, This represents the scene semantic classification weight matrix. This represents the bias vector. Indicates a category index; The complex scene recognition conclusion acquisition unit is used to aggregate the fault probability distribution and the scene semantic label to generate a diagnostic decision vector, parse the diagnostic decision vector to determine the fault diagnosis output, and obtain the complex scene recognition conclusion. The following formula is used to generate a diagnostic decision vector by weighted aggregation of the fault probability distribution and scene semantic labels: ; in, Represents the diagnostic decision vector. This represents the failure probability weight matrix. Represents the scene semantic weight matrix; The conclusion for complex scene recognition is derived using the following formula: ; in, This indicates the conclusion of complex scene recognition. This indicates the fault diagnosis output. Indicates the decision threshold. Indicates an indicator function; Fault diagnosis output This can be derived from the following formula: ; in, The first vector representing the diagnostic decision vector Quantity, This represents the fault category index.

2. The 5G MiFi device integrating voice and visual recognition according to claim 1, characterized in that, The semantic result acquisition module (10) includes: The instruction energy feature extraction unit is used to acquire a sampling sequence obtained by an audio acquisition device and converted, and to extract instruction energy features from the sampling sequence. The keyword extraction unit is used to encode the instruction energy features using a deep neural network model to generate an intent vector, and to extract search keywords based on the intent vector. The semantic result output unit is used to locate the parameter index or generate the optimization strategy based on the search keywords, and output semantic results related to network parameter query or signal optimization based on the parameter index or the optimization strategy.

3. The 5G MiFi device integrating voice and visual recognition according to claim 1, characterized in that, The optimization suggestion feedback data acquisition module (20) includes: The instruction type label determination unit is used to parse the intent slot information contained in the semantic result and compare the intent slot information with a preset instruction set to determine the instruction type label. An encrypted request message encapsulation unit is used to read the underlying registers to generate the original network status data and encapsulate it into an encrypted request message if the instruction type label is a network parameter query category. The network performance bottleneck feature acquisition unit is used to send the encrypted request message to the cloud server, and the cloud server inputs a large model to obtain the network performance bottleneck features. The optimization suggestion feedback data generation unit is used to generate optimization suggestion feedback data containing configuration adjustment parameters based on the network performance bottleneck characteristics.

4. The 5G MiFi device integrating voice and visual recognition according to claim 1, characterized in that, The preliminary assessment result acquisition module (30) includes: The spatiotemporal synchronization data packet generation unit is used to receive optimization suggestion feedback data sent from the cloud and visual acquisition data collected by the vision module, and to align the optimization suggestion feedback data and the visual acquisition data to generate a spatiotemporal synchronization data packet. The user gesture command encoding acquisition unit is used to extract image frames from the spatiotemporal synchronization data packet and calculate the user gesture command encoding. A legitimate control request signal generation unit is used to generate a legitimate control request signal if the user gesture instruction encoding matches the preset permission control requirements. The preliminary evaluation result output unit is used to construct a scene topology diagram based on the legitimate control request signal, and map the optimization suggestion feedback data to the scene topology diagram to output the preliminary evaluation result of multi-user network access permissions.

5. The 5G MiFi device integrating voice and visual recognition according to claim 1, characterized in that, The device connection list acquisition module (40) includes: The device access demand load calculation unit is used to parse the preliminary assessment results of multi-user network access permissions and calculate the device access demand load, which is determined based on the number of terminal requests and the service bandwidth demand value. The available spectrum resource block filtering unit is used to activate the wireless connection extension mechanism to filter available spectrum resource blocks if the device access demand load is greater than the single node access capacity threshold. The available spectrum resource blocks are extracted from the spare radio frequency channel. A multi-link concurrent access configuration table generation unit is used to generate a multi-link concurrent access configuration table based on the available spectrum resource block. The multi-link concurrent access configuration table is used to perform device registration and link migration. The device connection list acquisition unit is used to update the mapping relationship based on the multi-link concurrent access configuration table to obtain a dynamically adjusted device connection list.

Citation Information

Patent Citations

  • Edge cloud collaborative car networking multi-modal data analysis method

    CN118411748A

  • Software and hardware collaborative optimization method and device of multi-mode sensing and communication system

    CN119966820A