5g ai router integrating voice dialogue and screen display

By using microphone array beamforming and neural network model analysis, combined with distributed networks and clustering algorithms, the problem of synchronization between far-field voice commands and touch screens under 5G high-concurrency networks was solved, enabling real-time updates and precise control of device status, and improving network stability and user interaction efficiency.

CN121418345BActive Publication Date: 2026-03-20HUNAN YOUPIN IOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202512019112.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-20
Estimated Expiration
2045-12-30

AI Technical Summary

Technical Problem

In 5G high-concurrency dynamic networks, there is an inherent conflict between the synchronization of far-field voice commands and real-time display on the touch screen. This leads to real-time changes in device status data, resulting in inaccurate positioning, system misoperation or feedback of error information, and network chaos.

Method used

Microphone array beamforming technology is used to filter noise interference, and an operation intent vector is generated by parsing through a neural network model. Dynamic parameters are pulled from a distributed network in real time to update the device list. Clustering algorithm is used to group and sort the devices to determine the device location index, and a synchronous execution command sequence is generated through a real-time synchronization mechanism.

Benefits of technology

It achieves efficient and accurate matching and control of voice commands and dynamic devices, improves the robustness and response speed of far-field voice interaction, and ensures seamless execution of device operation in complex network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121418345B_ABST
    Figure CN121418345B_ABST
Patent Text Reader

Abstract

The application discloses a 5G AI router integrating voice dialogue and screen display, obtains a pure voice signal sequence by filtering noise through a microphone array beam forming technology, and adopts a neural network model to analyze and generate an accurate operation intention vector; a dynamic parameter fusion update device list structure is pulled from a distributed network in real time; when the intention vector target device identification and the list position do not match, a clustering algorithm is used to intelligently group and sort the list, and the matching device position index is quickly determined; then, a control parameter is extracted to align the voice requirements through a real-time synchronization mechanism, a synchronous execution command sequence is generated, the user interface is updated, and a network instruction is issued to adjust the device flow, and finally, the network state feedback data after execution is obtained. The application realizes efficient and accurate matching and control of voice instructions and dynamic devices, improves the robustness and response speed of far-field voice interaction, and ensures seamless execution of device operation in a complex network environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of routers, and discloses a 5G AI router integrating voice conversation and screen display. BACKGROUND

[0002] Under the background of popularization of high-speed networks in the 5G era, intelligent networking devices have become the core infrastructure for multi-terminal connection of families and small and micro enterprises, and their importance lies in directly affecting network stability and user daily operation convenience.

[0003] Current 5G router solutions are mostly limited to single wireless access or APP remote management, ignoring the complex requirements of real-time interaction in multi-device concurrent scenarios, which makes it difficult for users to quickly respond and adjust when the network is congested or abnormal.

[0004] This limitation is further magnified on integrated devices that combine voice and screen, because the 5G dual-mode high-concurrent connection feature causes the device state data to change dramatically in real time, and there is a natural conflict between the recognition range of far-field voice instructions and the synchronization of touch screen display: the voice array needs to cover multiple directions within 5 meters to pick up sound to capture user instructions, but dynamic parameters such as speed and traffic statistics under high-concurrent networks change rapidly, if the voice analysis cannot accurately match the real-time visual position on the screen, there will be instruction execution deviation, for example, when the user says "limit speed child device", the device position in the screen list has shifted due to new analysis-connection refresh, causing the system to misoperate other devices or feedback incorrect information, thereby causing network chaos in a multi-user home environment.

[0005] Therefore, how to realize seamless and accurate synchronization between far-field voice instructions and real-time display of touch screens under the 5G high-concurrent dynamic network has become a key problem for building integrated intelligent routers. SUMMARY

[0006] The application provides a 5G AI router integrating voice conversation and screen display, which aims to solve at least one of the defects in the prior art.

[0007] The application relates to a 5G AI router integrating voice conversation and screen display, comprising:

[0008] An instruction content determination module is configured to capture voice instructions issued by a user from a far-field voice pickup array, obtain a preliminary voice signal sequence through microphone array signal processing, enhance directional pickup to filter noise interference by using a beam forming technology, and determine instruction content containing keywords.

[0009] The operation intention vector generation module is configured to perform semantic analysis on the speech signal sequence according to the instruction content by using a neural network model, and generate a corresponding operation intention vector, the operation intention vector representing an identifier of the target device and a parameter requirement;

[0010] The list structure acquisition module is configured to acquire current device list data on the user interface device, pull dynamic parameter monitoring information from the distributed network in real time, fuse the dynamic parameter monitoring information into the current device list data to update device states, and obtain a refreshed list structure.

[0011] The device position index determination module is configured to, if the target device identifier in the operation intention vector does not match a position in the list structure, group and sort the list structure by using a clustering algorithm, and determine a matched device position index.

[0012] The synchronous execution command generation module is configured to extract a control parameter from the device position index, align the control parameter with a requirement of the voice instruction by using a real-time synchronization mechanism, and generate a synchronous execution command sequence.

[0013] The network state feedback data acquisition module is configured to update content of the user interface device by using the synchronous execution command sequence, and simultaneously issue an instruction to the distributed network to adjust a target device flow, and obtain network state feedback data after execution.

[0014] Further, the 5G AI router integrating voice dialogue and screen display provided in the embodiment includes an instruction content determination module.

[0015] The preliminary speech signal sequence acquisition unit is configured to acquire a far-field speech signal sequence captured by a microphone array, the far-field speech signal sequence being a voice instruction issued by a user, and obtain a preliminary speech signal sequence.

[0016] The directional speech signal sequence acquisition unit is configured to, if signal energy of the preliminary speech signal sequence exceeds a preset threshold, determine a directional enhancement signal by using delay-sum beamforming processing on the preliminary speech signal sequence, and obtain a directional speech signal sequence.

[0017] The preliminary speech feature sequence acquisition unit is configured to acquire a frequency spectrum feature and a time domain feature from the directional speech signal sequence, and obtain a preliminary speech feature sequence.

[0018] The instruction content determination unit is configured to, if the preliminary speech feature sequence contains a keyword, determine an instruction content containing the keyword by using an end-to-end speech recognition model on the preliminary speech feature sequence, and determine the instruction content.

[0019] Further, the operation intention vector generation module includes:

[0020] The semantic parsing sequence acquisition unit is configured to perform semantic parsing on the speech signal sequence by using a convolutional neural network according to the instruction content, to obtain a semantic parsing sequence.

[0021] The target device identifier sequence acquisition unit is configured to acquire a target device identifier from the semantic parsing sequence if the confidence of the semantic parsing sequence exceeds a preset threshold, to obtain a target device identifier sequence.

[0022] The parameter requirement vector determination unit is configured to match the target device identifier sequence with a pre-established parameter requirement database, to determine a parameter requirement vector.

[0023] The operation intention vector acquisition unit is configured to fuse the target device identifier sequence and the parameter requirement vector by using a vector splicing operation if the parameter requirement vector matches the target device identifier sequence successfully, to obtain an operation intention vector.

[0024] Further, the list structure acquisition module comprises:

[0025] The initial device list acquisition unit is configured to acquire current device list data in the user interface device, to extract current device identifiers and basic attributes of the current device list data by using a database query operation, and to obtain an initial device list.

[0026] The dynamic parameter data set acquisition unit is configured to pull dynamic parameter monitoring information corresponding to the current device identifiers in real time through a distributed network interface if the initial device list contains at least one current device identifier, to obtain a dynamic parameter data set.

[0027] The updated device list acquisition unit is configured to match and integrate the dynamic parameters and the current device identifiers in the initial device list by using a data fusion algorithm according to the dynamic parameter data set, to obtain an updated device list.

[0028] The list structure acquisition unit is configured to load the updated device list to the user interface device by using an interface refreshing operation if the updated device list conforms to a preset list structure format, to obtain a refreshed list structure.

[0029] Further, the device position index determination module comprises:

[0030] The preliminary matching result acquisition unit is configured to acquire a target device identifier in the operation intention vector and a device identifier set in the list structure, to determine whether the target device identifier is identical to any current device identifier in the list structure by using a comparison operation, and to obtain a preliminary matching result.

[0031] The device identification sequence acquisition unit is configured to, if the preliminary matching result indicates that the target device identification is inconsistent with the current device identification in the list structure, group and sort the current device identification in the list structure based on dynamic parameters through a clustering algorithm to obtain a sorted device identification sequence.

[0032] The device position index determination unit is configured to compare the target device identification with the current device identification in the sorted device identification sequence one by one through an index mapping operation to determine a matched device position index.

[0033] Further, the synchronous execution command generation module comprises:

[0034] The preliminary synchronization parameter sequence determination unit is configured to acquire a device control parameter set from the device position index, compare the control parameter set with parameter requirements in the voice instruction one by one through a real-time synchronization mechanism to determine a preliminary synchronization parameter sequence.

[0035] The synchronization parameter sequence acquisition unit is configured to, if there is a deviation between the parameter value in the preliminary synchronization parameter sequence and the requirement of the voice instruction, determine whether the deviation meets the requirement through a preset threshold to obtain an adjusted synchronization parameter sequence.

[0036] The synchronous execution command sequence generation unit is configured to, according to the adjusted synchronization parameter sequence, convert the parameter sequence into an executable command sequence through a command generation algorithm to generate a synchronous execution command sequence.

[0037] Further, the network state feedback data acquisition module comprises:

[0038] The display content update unit is configured to update the content displayed on the user interface device according to the synchronous execution command sequence.

[0039] The network state feedback data acquisition unit is configured to issue an instruction to the distributed network to adjust the target device flow to obtain executed network state feedback data.

[0040] The present application has the following beneficial effects:

[0041] The application discloses a 5G AI router integrating voice dialogue and screen display, and aims at a service scene problem that a target device position in a dynamic device list does not match a user voice instruction, wherein the problem is caused by noise interference of voice pickup, semantic analysis deviation and positioning deviation caused by real-time change of a distributed network device state. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 It is a functional block diagram of one embodiment of the 5G AI router integrating voice dialogue and screen display.

[0043] REFERENCE SIGNS

[0044] 10, instruction content determination module; 20, operation intention vector generation module; 30, list structure acquisition module; 40, device position index determination module; 50, synchronous execution command generation module; 60, network state feedback data acquisition module. DETAILED DESCRIPTION

[0045] In order to better understand the above technical solutions, the above technical solutions will be described in detail in combination with the drawings in the specification and specific embodiments.

[0046] As Figure 1As shown, the first embodiment of the present application proposes a 5G AI router integrating voice dialogue and screen display, which comprises an instruction content determination module 10, an operation intention vector generation module 20, a list structure acquisition module 30, a device position index determination module 40, a synchronous execution command generation module 50 and a network state feedback data acquisition module 60. The instruction content determination module 10 is used to capture the voice instruction issued by the user from the far-field voice pickup array, obtain the preliminary voice signal sequence through the microphone array signal processing, enhance the directivity pickup by using the beam forming technology to filter the noise interference, and determine the instruction content containing the keyword. The operation intention vector generation module 20 is used to perform semantic analysis on the voice signal sequence by using the neural network model according to the instruction content, and generate the corresponding operation intention vector. The operation intention vector represents the identification and parameter requirement of the target device. The list structure acquisition module 30 is used to acquire the current device list data on the user interface device, pull the dynamic parameter monitoring information from the distributed network in real time, fuse into the current device list data to update the device state, and obtain the refreshed list structure. The device position index determination module 40 is used to group and sort the list structure by using the clustering algorithm if the target device identification in the operation intention vector does not match the position in the list structure, and determine the matched device position index. The synchronous execution command generation module 50 is used to extract the control parameter from the device position index, align the control parameter with the requirement of the voice instruction by using the real-time synchronization mechanism, and generate the synchronous execution command sequence. The network state feedback data acquisition module 60 is used to update the user interface device content by using the synchronous execution command sequence, simultaneously issue the instruction to the distributed network to adjust the target device flow, and obtain the executed network state feedback data.

[0047] The instruction content determination module 10 is used to capture the voice instruction issued by the user in an open environment such as home or office by using a far-field pickup array composed of 6-8 omnidirectional microphones (covering a 360° pickup range, effective distance 3-8 meters) in synchronization; the multi-channel voice signal is preprocessed by a microphone array signal processing technology (such as time delay estimation, signal alignment) to generate a time-synchronized preliminary voice signal sequence; then an adaptive beam forming technology (such as MVDR (Minimum Variance Distortionless Response) algorithm) is used to construct a directional pickup beam, focus on the user sound source direction (angle error ≤5°), and suppress the environmental noise (such as TV sound, footstep sound) and echo interference (noise reduction amount ≥25dB) in the non-target direction; finally, through endpoint detection and keyword spotting algorithm, the keyword combination containing “target device identifier” (such as “living room TV”), “operation action” (such as “limit speed”), and “parameter requirement” (such as “100 Mbps”) is extracted from the enhanced voice signal to form a structured instruction content (such as “living room TV limit speed 100 Mbps”), ensuring that the instruction information is complete and the noise interference rate is ≤5%.

[0048] The operation intention vector generation module 20 is used to process the voice instruction containing keywords (such as “living room TV limit speed to 50 Mbps”) output by the instruction content determination module 10. First, the voice signal sequence is converted into features (extracting Mel-frequency cepstral coefficients MFCC (Mel-Frequency Cepstral Coefficients)), and then input into a pre-trained BERT-BiLSTM neural network model for semantic analysis. The BERT-BiLSTM neural network model identifies the core elements in the instruction through context association analysis: target device unique identifier (such as MAC (Media Access Control Address) address hash value, device alias code), operation type (such as bandwidth adjustment, connection control, state query), and specific parameter requirement (such as rate value, time threshold, switch state). These elements are quantified into 128-dimensional operation intention vectors, where the first 64 dimensions are used to encode device identifiers (supporting 264 unique device identifiers), and the last 64 dimensions are used to describe parameter requirements (including operation type encoding, parameter value, and unit conversion coefficient). The vector cosine similarity matching accuracy is ≥98%, ensuring that the intention analysis is unambiguous, and providing standardized input for device positioning and command generation.

[0049] The list structure acquisition module 30 is used to acquire the basic data of the current online device list (including device name, MAC address, IP address, connection type, etc. static information) from the router local user interface (screen display / APP end); through the edge nodes (such as sub-routers, IoT (Internet of Things) gateways) in the distributed network, the dynamic parameter monitoring information (updated every 3 seconds) of the device is pulled in real time by using the MQTT (Message Quueuing Telemetry Transport) protocol, including real-time upload / download rate, connection time, signal strength, traffic usage, etc.; the dynamic parameters and the local list data are associated and matched according to the MAC address by using a data fusion algorithm, the device state field (such as "online / offline", "speed limiting") is updated, data conflicts are handled (such as the rate value deviation between the local and the edge node is greater than 10%, the weighted average value is taken), and finally a refreshed list structure containing 20+ fields is generated, which supports multi-dimensional sorting according to "device type", "room partition" and "network load", ensures the real-time performance (delay≤500ms) and accuracy (field matching error<1%) of the list data.

[0050] The device position index determination module 40 is used to when the target device identification (such as MAC hash value, alias code) in the operation intention vector cannot be directly matched with the device position index (such as the 5th item in the list) in the refreshed list structure (matching failure rate>5%), the multi-dimensional features (type, position, connection time, signal strength, etc.) of the device in the list structure are extracted, the device is grouped (such as "mobile device group", "living room IoT group") by using an improved K-means clustering algorithm; based on "device identification similarity+usage frequency" within the group, the target device identification and the matching score of the devices in the group (such as MAC address fuzzy matching degree, alias semantic similarity) are calculated; the device position with the highest score is selected as the matching index, if the score is greater than or equal to 80, the index is directly determined, otherwise the secondary clustering is triggered (the "historical control record" feature is added); finally, the matched device position index (such as the 2nd item in the 3rd group) is output, ensuring that the positioning accuracy is greater than or equal to 98% and the response time is less than or equal to 200ms.

[0051] The synchronous execution command generation module 50 is configured to extract the current control parameters (such as real-time speed, connection priority, bandwidth limit threshold) of the target device from the list item corresponding to the device location index; based on the parameter requirements (such as "limit speed to 50 Mbps") in the operation intention vector, the control parameters are dynamically aligned with the instruction requirements by using a timestamp synchronization mechanism, the parameter adjustment amount (such as the current 80 Mbps needs to be reduced by 30 Mbps) is calculated by using a deviation compensation algorithm, the command unit containing the timestamp, the instruction type, the parameter value, and the check code is generated in combination with the device response characteristics (such as the delay compensation coefficient of the IoT device), the command unit is combined into a synchronous execution command sequence (supporting single-device multi-parameter linkage or multi-device concurrent control) according to the execution time sequence, the time interval of adjacent commands in the sequence is greater than or equal to the minimum response period of the device (such as 100 ms), and it is ensured that the next command is triggered after the previous command is executed; and finally, the sequence needs to pass the syntax check (in line with the device control protocol) and the safety check (the parameters are within the safety threshold), and the execution success rate is greater than or equal to 99%.

[0052] The network state feedback data acquisition module 60 is configured to split the synchronous execution command sequence into "interface update instructions" and "traffic regulation instructions" for issuing, the interface update instructions are pushed to the local screen of the router and the associated APP in real time, and the parameter display in the user interface device list is updated (such as changing the network speed of the "living room TV" from 80 Mbps to 50 Mbps, and marking "adjusted" in green); the traffic regulation instructions are issued to the main router traffic control module and the edge nodes (sub-routers, IoT gateways) through a distributed network protocol (such as OpenFlow, MQTT), and the bandwidth allocation, traffic priority or connection permission of the target device is adjusted; after the execution of the instructions, the network state data (sampling frequency 1 Hz) including the actual speed, network delay, packet loss rate, traffic usage and interface update state of the target device are collected in real time through the SNMP protocol and the device end probe; the collected data is cleaned (abnormal values are removed) and integrated (aligned according to the timestamp), and the feedback data report containing "execution result-parameter change-network stability" is generated, the feedback delay is less than or equal to 1 second, the data accuracy is greater than or equal to 98%, and it is ensured that the user can know the control effect in real time through the interface or voice broadcast.

[0053] Further, the 5G AI router integrating voice dialogue and screen display provided by the embodiment includes an initial voice signal sequence acquisition unit, a directional voice signal sequence acquisition unit, an initial voice feature sequence acquisition unit, and an instruction content determination unit. The initial voice signal sequence acquisition unit is configured to acquire a far-field voice signal sequence captured by a microphone array. The far-field voice signal sequence is a voice instruction issued by a user, and an initial voice signal sequence is obtained.

[0054] The following formula describes how to obtain a preliminary speech signal sequence by weighting and superimposing signals captured by multiple microphones and correcting for time delay:

[0055] (1)

[0056] In formula (1), This represents the initial speech signal sequence after preprocessing. This indicates the total number of microphones in the microphone array. Indicates the first Weighting coefficients for each microphone, Indicates the first Each microphone at any time Captured far-field speech signals, Indicates the first The delay compensation parameters for each microphone. The control logic of formula (1) is to first "synchronize" the signals of each microphone to the same time point, and then "mix them together" according to their importance, so as to obtain a clearer preliminary speech sequence.

[0057] A microphone array consists of multiple microphones capable of capturing sound signals from different directions. These signals are then converted into a sequence by a digital signal processor to obtain a preliminary speech signal sequence. This process involves sampling the analog speech signal into a digital sequence at a sampling rate of perhaps 16 kHz to ensure sufficient speech detail is captured, thus providing the foundational data for subsequent processing.

[0058] Delay-sum beamforming is applied to the initial speech signal sequence. This method is a spatial filtering technique designed to enhance speech signals from a specific direction and suppress noise. In principle, it calculates the time delay between each microphone signal, sums the signals from the target direction, and thus forms a beam. For example, in an array of four microphones, if the user's voice comes from the front, the system estimates the delay time, such as 0.1 milliseconds for the first microphone signal and 0.2 milliseconds for the second. These signals are then phase-aligned and summed to obtain the enhanced signal.

[0059] The directional speech signal sequence acquisition unit is used to perform delayed summation beamforming processing on the preliminary speech signal sequence. If the signal energy of the preliminary speech signal sequence exceeds a preset threshold, the directional enhancement signal is determined, and the directional speech signal sequence is obtained.

[0060] The signal energy of the initial speech signal sequence is obtained using the following formula:

[0061] (2)

[0062] In formula (2), This represents the signal energy of the initial speech signal sequence. Indicates the total length of the signal sequence. Indicates the first The signal amplitude at each sampling point is used to calculate the average of the squared amplitudes of all sampling points to obtain the total energy of the signal. The control logic of formula (2) is to add up the "intensity contribution" (squared amplitude) of each point of the signal and then average it to each sampling point to measure the energy of the entire signal.

[0063] The directional speech signal sequence is obtained using the following formula:

[0064] (3)

[0065] In formula (3), The first directional speech signal sequence One frequency component, The gain factor that represents the enhancement of directionality. The input signal is represented by the first... One frequency component, Indicates the first The power value of each frequency component This indicates the preset energy judgment threshold. The control logic of formula (3) is to "select the effective speech frequency components that meet the energy standard and amplify them, and not to further enhance the interference components with weak energy", so as to highlight the speech signal in the target direction.

[0066] If the signal energy of the initial speech signal sequence exceeds a preset threshold, for example, a threshold set to -40dB, it is determined to be a directional enhancement signal. This means that the system has detected sufficient speech intensity to avoid processing background noise. This results in a directional speech signal sequence with a significantly improved signal-to-noise ratio, which is beneficial for subsequent feature extraction.

[0067] The preliminary speech feature sequence acquisition unit is used to acquire spectral and temporal features from the directional speech signal sequence to obtain the preliminary speech feature sequence.

[0068] The following formula is used to convert the directional speech signal in the time domain into a frequency domain representation using the Discrete Fourier Transform, for extracting spectral features:

[0069] (4)

[0070] In formula (4), Frequency domain representation of directional speech signals This represents the input directional speech signal sequence. Represents the window function. Indicates signal length. Indicates the frequency bin index, The base function represents a discrete Fourier transform. The control logic of formula (4) is to "window and optimize the time-domain signal first, and then convert it into a frequency-domain frequency component representation through a DFT (Discrete Fourier Transform) operation, to facilitate subsequent speech processing (such as noise reduction and enhancement) in the frequency domain.

[0071] The short-time energy of each frame of speech signal is calculated by the following formula, which is an important time-domain feature for feature extraction of the speech signal:

[0072] (5)

[0073] In formula (5), the time-domain energy feature of the i-th frame is represented by , the frame length is represented by , the frame shift is represented by , the frame index is represented by , the original speech signal is represented by , and the square of the modulo operation on the signal is represented by The control logic of formula (5) is "cut the long speech into short frames, and then calculate the average energy of each small segment" to characterize the energy changes of the speech signal at different time periods.

[0074] The extracted spectral features and time-domain features are fused by weighting to form the final preliminary speech feature sequence by the following formula:

[0075] (6)

[0076] In formula (6), the preliminary speech feature vector at the i-th time is represented by , the spectral feature vector is represented by , the time-domain feature vector is represented by , the weight coefficient of the spectral feature is represented by , and the weight coefficient of the time-domain feature is represented by The control logic of formula (6) is "assign values to the spectral and time-domain features according to their importance, and then combine them into a comprehensive feature", so that the final feature covers speech information in different dimensions.

[0077] Spectral features and time-domain features are obtained from the sequence of directional speech signals to obtain a preliminary speech feature sequence. The spectral features can be converted from the time domain to the frequency domain by Fourier transform to extract, for example, Mel-Frequency Cepstral Coefficients (MFCC), which simulate human perception of sound, such as calculating the energy distribution of the signal on different frequency bands. The time-domain features include zero-crossing rate or short-time energy, which are used to capture the time-varying characteristics of the signal. The specific process is to first frame the sequence, 20 milliseconds per frame, then apply a Hamming window function for smoothing, calculate the spectrum and extract the feature vector, and finally combine the preliminary speech feature sequence.

[0078] The instruction content determination unit is configured to process the preliminary speech feature sequence using an end-to-end speech recognition model, and if the preliminary speech feature sequence contains a keyword, determine the instruction content containing the keyword.

[0079] The determined instruction content is obtained by the following formula:

[0080] (7)

[0081] In formula (7), represents the determined instruction content, represents the candidate instruction category, represents the set of all possible instruction categories, represents the number of recognized keywords, represents the importance weight of the th keyword, represents the feature of the th keyword, The function represents a matching degree evaluation function of the keyword and the instruction category. The control logic of formula (7) is "give each keyword a score according to its importance, calculate their total score with each candidate instruction, and select the instruction with the highest score as the result", which realizes accurate mapping from keywords to instructions.

[0082] The output result of the end-to-end speech recognition model is obtained by the following formula:

[0083] (8)

[0084] In formula (8), represents the output result of the end-to-end speech recognition model, represents the input preliminary speech feature sequence, represents the end-to-end speech recognition model, represents the encoder part for extracting deep features, The decoder part is used to generate the recognition result. The control logic of formula (8) is that "the encoder compresses and refines the speech features, and the decoder translates the refined information into the recognition result", realizing the direct mapping from speech features to the final output.

[0085] The initial speech feature sequence (such as Mel spectrogram and MFCC features) obtained after microphone array processing is used as input and fed into a pre-trained end-to-end speech recognition model (such as Conformer-Transducer). The end-to-end speech recognition model performs deep modeling of temporal features through encoder, captures semantic association information in speech, and then directly maps the feature sequence into a text sequence through decoder. At the same time, the end-to-end speech recognition model has a built-in keyword detection branch (focusing on a preset keyword library such as "device name" and "operation verb" based on attention mechanism) to determine whether the text sequence contains keywords in real time. If keywords are detected (such as "living room TV" and "speed limit"), keyword combinations and associated semantics (such as "50Mbps") are extracted from the text sequence to form a structured instruction content containing "target device - operation action - parameter requirements". If no keywords are detected, an "invalid instruction" prompt is returned, triggering secondary sound pickup (such as "please repeat instruction"). The entire process has a response time of ≤500ms, a keyword recognition accuracy of ≥96%, and an instruction content completeness of ≥95%.

[0086] Furthermore, the 5G AI router integrating voice dialogue and screen display provided in this embodiment includes an operation intent vector generation module 20 comprising a semantic parsing sequence acquisition unit, a target device identifier sequence acquisition unit, a parameter requirement vector determination unit, and an operation intent vector acquisition unit. The semantic parsing sequence acquisition unit is used to perform semantic parsing on the speech signal sequence using a convolutional neural network according to the instruction content to obtain a semantic parsing sequence.

[0087] The semantic parsing sequence is derived using the following formula:

[0088] (9)

[0089] In formula (9), Indicates the first The final semantic parsing result at each time step, Represents the set of all possible semantic categories. This represents the total number of convolutional feature maps. Indicates the first The classification weight vector corresponding to each feature map Indicates the first Convolutional feature vectors at each time step Indicates the first Threshold parameters for each classifier is the activation function. The control logic of formula (9) is to "let multiple convolutional feature maps score the semantic category respectively, take the average score and select the category with the highest score", thereby improving the reliability of semantic parsing through the fusion of multiple feature maps.

[0090] Based on the instruction, a convolutional neural network (CNN) is used to perform semantic parsing on the speech signal sequence, resulting in a semantically parsed sequence. A CNN is a deep learning model that extracts features through multiple convolutional and pooling layers, making it particularly suitable for processing sequential data such as speech signals. The process first involves inputting the speech signal sequence into the network's input layer. Then, convolutional operations capture local patterns, such as the convolutional kernel sliding through the sequence to identify semantic patterns, like verbs and nouns in the instruction. Next, the hidden layers of the CNN gradually abstract these features, and finally, the output layer generates the semantically parsed sequence. This sequence may include parsed intent labels and entity recognition results. For example, in a smart home scenario, if a user says "adjust the living room light to 50% brightness," the speech signal sequence, after processing by the CNN, yields a semantically parsed sequence, where "adjust" is identified as the action intent, "living room light" as the device entity, and "50% brightness" as the parameter value, thus forming a structured sequence representation.

[0091] The target device identifier sequence acquisition unit is used to obtain the target device identifier from the semantic parsing sequence if the confidence level of the semantic parsing sequence exceeds a preset threshold, thereby obtaining the target device identifier sequence.

[0092] The confidence score of the semantically parsed sequence is calculated using the following formula:

[0093] (10)

[0094] In formula (10), This represents the confidence level of the semantic parsing sequence. Indicates the length of the parsed sequence. Indicates the condition that all preceding words are given. The probability of each word is used to evaluate the credibility of the entire semantic parsing sequence by calculating the average conditional probability of all words in the sequence. The control logic of formula (10) is to see how smoothly each word is connected in the current sequence, and to take the average of the smoothness of all words to judge whether the entire parsing result is reliable.

[0095] The following formula is used to define the confidence threshold for decision-making:

[0096] (11)

[0097] In formula (11), This indicates the threshold judgment decision result. This represents the confidence value of the current semantic parsing sequence. represents a preset confidence threshold, when the confidence exceeds the threshold, 1 is returned to execute the subsequent extraction operation, otherwise 0 is returned without execution. The control logic of formula (11) is “use the threshold as the passing line, continue if the confidence exceeds the line, and terminate if it does not”, so as to filter the unreliable semantic analysis results.

[0098] The target device identification sequence is obtained by the following formula:

[0099] (12)

[0100] In formula (12), represents the extracted target device identification sequence, containing device identifications to , represents the input semantic analysis sequence, represents the pattern matching rule for identifying and extracting device identifications, and all target device identifications are extracted from the semantic sequence by the Extract function. The control logic of formula (12) is “use the preset device identification template to find the corresponding content in the semantic sequence, and arrange the found results into a device identification sequence”, which realizes the extraction from semantic information to target device identification.

[0101] If the confidence of the semantic analysis sequence exceeds the preset threshold, the target device identification is obtained from the semantic analysis sequence, and the target device identification sequence is obtained. The confidence is a quantitative of the reliability of the analysis result by the deep learning model, and the probability value is usually calculated by the softmax function, and the preset threshold is, for example, 0.8, which represents high confidence.

[0102] Specifically, the confidence score of each element in the semantic analysis sequence is checked, and if the overall average exceeds the threshold, the identification part is extracted. For example, the confidence of the semantic analysis sequence showing the “open navigation” intent is 0.9, which exceeds the 0.7 threshold, so “navigation” is extracted from the semantic analysis sequence as the target device identification, forming an identification sequence such as [navigation, screen], which ensures the accurate basis for subsequent operations.

[0103] The parameter requirement vector determination unit is configured to match the parameter requirement database established in advance according to the target device identification sequence, and determine the parameter requirement vector.

[0104] The parameter requirement vector is obtained by the following formula:

[0105] (13)

[0106] In formula (13), represents the determined parameter requirement vector value, represents the dimension number of the parameter requirement, importance coefficient of the base value of the function represents a step function, confidence of the represents a confidence threshold. The control logic of formula (13) is "first filter out the untrusted parameters, then add up the importance of the remaining trusted parameters", which guarantees the reliability and rationality of the parameter requirement vector.

[0107] According to the target device identification sequence, the parameter requirement database is matched to determine the parameter requirement vector. The parameter requirement database is a pre-constructed storage structure, which contains the mapping of device identification and corresponding parameters, such as key-value pair storage. Specifically, the matching process involves querying the database, using the identification sequence as the key to find the related parameters, for example, for the "air conditioner" identification, the parameter requirement database returns the vector [temperature range: 18-30, mode: cooling / heating], so as to determine the parameter requirement vector. This vector is a numerical representation, which is convenient for further processing.

[0108] The operation intention vector acquisition unit is used to fuse the target device identification sequence and the parameter requirement vector by vector splicing operation if the parameter requirement vector and the target device identification sequence match successfully, to obtain the operation intention vector.

[0109] The matching function of the parameter requirement vector and the target device identification sequence is defined by the following formula:

[0110] (14)

[0111] In formula (14), represents the matching function of the parameter requirement vector and the target device identification sequence, represents the parameter requirement vector, represents the target device identification sequence, represents the matching threshold, and returns 1 when the Euclidean distance of the two vectors is less than or equal to the threshold, otherwise returns 0. The control logic of formula (14) is "calculate the distance between the two vectors, if the distance is within the threshold, it is considered as a match, otherwise it is not considered as a match", which realizes the matching verification of the parameter requirement and the target device.

[0112] The operation intention vector is obtained by the following formula:

[0113] (15)

[0114] In formula (15), represents the operation intention vector, and represents the vector splicing operation,​​​ represents the i-th element of the target device identification sequence, represents the i-th element of the parameter requirement vector, represents the i-th element of the parameter requirement vector, represents the i-th element of the parameter requirement vector, and respectively represent the dimensions of two vectors. The control logic of formula (15) is to splice the device to be operated and the parameter requirement of the operation together to form a complete operation intention description.

[0115] If the parameter requirement vector matches the target device identification sequence successfully, the target device identification sequence and the parameter requirement vector are fused by vector splicing operation to obtain an operation intention vector. Successful matching means that the parameters in the vector are consistent with the identification and there is no conflict. Specifically, vector splicing is to connect two vectors along the dimension, for example, splicing the identification sequence [light, living room] and the parameter vector [brightness: 50, color: warm white] into [light, living room, brightness: 50, color: warm white] to form an operation intention vector. For example, the user instruction "set the conference room projector to high-definition mode" is matched and spliced to obtain an intention vector, which is used to directly drive device control, which improves the accuracy of instruction execution in the business.

[0116] Further, the 5G AI router integrating voice dialogue and screen display provided by the embodiment provides a list structure acquisition module 30, which includes an initial device list acquisition unit, a dynamic parameter data set acquisition unit, an updated device list acquisition unit, and a list structure acquisition unit. The initial device list acquisition unit is configured to acquire current device list data in the user interface device, extract current device identifiers and basic attributes of the current device list data by using a database query operation, and obtain an initial device list.

[0117] The process of forming a device list by extracting device identifiers and attributes by a database query operation is described by the following formula:

[0118] (16)

[0119] In formula (16), represents a set of initial device list data obtained from a database query, represents the total number of devices detected in the current user interface, represents the i-th device identifier, represents the i-th device identifier, represents the i-th device identifier, represents the i-th device identifier,

[0120] When obtaining the current device list data in the user interface device, first, key information is extracted through a database query operation. Specifically, this database query operation is a retrieval mechanism based on the SQL (Structured Query Language) language, which allows the system to pull out the current device identification, such as the device ID number, and basic attributes, such as the device type and connection state, from the table storing the device information, thereby forming an initial device list. For example, the user interface device such as a smart phone APP will query the local SQLite database to extract the identification "Light_001" of the living room lamp and its basic attributes "type: LED, status: online", and integrate these information into a list structure, ensuring reliable basic data for subsequent steps. In this way, the initial device list not only records static information, but also provides an entry point for dynamic updating, which can improve the real-time performance of device management in business.

[0121] The dynamic parameter data set acquisition unit is configured to, if the initial device list contains at least one current device identification, pull real-time dynamic parameter monitoring information corresponding to the current device identification through a distributed network interface to obtain a dynamic parameter data set.

[0122] The dynamic parameter data set is obtained by the following formula:

[0123] (17)

[0124] In formula (17), represents the dynamic parameter data set obtained at time represents the total number of devices in the network, represents the weight coefficient of the th device, represents the parameter monitoring information of the th device at time represents an indicator function, represents the identification of the th device, represents the initial device list. The control logic of formula (17) is to first filter out the devices in the initial list, and then add up the real-time parameters of these devices according to the weight to obtain the dynamic parameter set at the current time.

[0125] ​​If the initial device list contains at least one current device identifier, the corresponding dynamic parameter monitoring information is retrieved in real time via a distributed network interface. The distributed network interface is a communication protocol based on RESTful API (Representational State Transfer), which allows the system to obtain real-time data from cloud servers or edge nodes, such as the device's current power consumption or temperature readings, forming a dynamic parameter dataset. For example, if the initial list contains the identifier "AC_002" for "air conditioner," the system will retrieve dynamic parameters from the vehicle's sensors via the distributed interface of the MQTT (Message Queuing Telemetry Transport) protocol, such as the current temperature of 25 degrees Celsius and wind speed level 3. This data is aggregated into a dataset [temperature: 25, wind speed: 3], ensuring the freshness of the information and avoiding operational errors caused by outdated data.

[0126] The device list acquisition unit is used to match and integrate the dynamic parameters with the current device identifier in the initial device list based on the dynamic parameter dataset using a data fusion algorithm, so as to obtain the updated device list.

[0127] The similarity score between device identification and dynamic parameters is calculated using the following formula:

[0128] (18)

[0129] In formula (18), Indicates the current device identifier With dynamic parameters Similarity matching score between them Indicates the first device in the initial device list A feature vector identifying the current device. Represents the first in the dynamic parameter dataset Feature vectors with parameters The bandwidth parameter of the Gaussian kernel function is used to control the sensitivity of the matching. The control logic of formula (18) is to measure the difference by distance, and then use the Gaussian function to convert the difference into a similarity score between 0 and 1. The bandwidth controls the sensitivity of the score to the difference.

[0130] According to the dynamic parameter dataset, a data fusion algorithm is used to match and integrate the dynamic parameters with the device identifiers in the initial device list. The data fusion algorithm is a multi-source information integration technology that associates dynamic data with static identifiers through key-value matching and weighted averaging methods to form an updated device list. For example, for the "sensor_003" identifier and its basic attributes "type: temperature sensor" in the initial list, the data fusion algorithm matches the real-time readings in the dynamic dataset, such as humidity 45% and voltage 3.2V, confirms consistency through identifier matching, and then integrates them into the updated list item [sensor_003, type: temperature sensor, humidity: 45%, voltage: 3.2V]. This process, from data cleaning to feature alignment, ensures the completeness and accuracy of the list. In business, this can bring more accurate device monitoring results, such as timely detection of abnormal parameters to prevent faults.

[0131] The list structure acquisition unit is configured to, if the updated device list conforms to the preset list structure format, load the updated device list to the user interface device through an interface refresh operation to obtain a refreshed list structure.

[0132] The verification condition of the updated device list is defined by the following formula:

[0133] (19)

[0134] In formula (19), represents the verification result of the updated device list, represents the updated device list, represents the structure format of the updated device list, represents the preset list structure format, and returns 1 when the structure format of the updated device list completely matches the preset format, otherwise returns 0. The control logic of formula (19) is to check whether the structure of the updated device list is correct, and if it is completely the same as the preset format, it passes the verification, otherwise it does not pass.

[0135] The process of loading the verified device list to the user interface through the interface refresh operation is described by the following formula:

[0136] (20)

[0137] In formula (20), represents the update operation of the user interface, represents the verified device list, represents the interface refresh function, The time parameter representing the refresh operation. The control logic of formula (20) is to use the list of devices passed the verification as the data source, trigger the interface refresh to load the latest device list according to the set time rhythm.

[0138] The final list structure obtained by the following formula is the result of the original list and the updated change amount:

[0139] (21)

[0140] In formula (21), represents the final list structure after refresh, represents the original list structure, represents the change amount of list update, and the symbol represents the list merge update operation. The control logic of formula (21) is to add the content to be updated to the original list to merge into the final new list.

[0141] If the updated device list meets the preset list structure format, it is loaded to the user interface device through the interface refresh operation. The interface refresh operation is an event-driven UI (User Interface) update mechanism, which checks whether the list meets the standard structure of the JSON format, such as each item has a fixed field, and then renders to the interface to obtain the refreshed list structure. For example, if the updated list meets the format of {device ID: attribute object}, the system will use the DOM (Document Object Model) operation of JavaScript to refresh the APP interface and display, for example, "monitor_004: online, heart rate: 80bpm", which improves the smoothness of user interaction.

[0142] Further, the 5G AI router integrating voice dialogue and screen display provided by the embodiment includes a preliminary matching result acquisition unit, a device identification sequence acquisition unit, and a device position index determination unit. The preliminary matching result acquisition unit is configured to acquire a target device identification in an operation intention vector and a device identification set in a list structure, and to determine whether the target device identification is consistent with any current device identification in the list structure by using a comparison operation to obtain a preliminary matching result.

[0143] The preliminary matching result is obtained by the following formula:

[0144] (22)

[0145] In formula (22), represents the preliminary matching result, represents the total number of devices in the device identification set, represents the target device identifier extracted from the operation intention vector, represents the i-th device identifier in the list structure, represents the i-th device identifier in the list structure, The function represents a comparison operation function, which determines whether there is a match by comparing all device identifiers and taking the maximum value. The control logic of formula (22) is to check one by one whether there is a device in the device list that is the same as the target device. If there is a match, the result is success, otherwise, it is failure.

[0146] The judgment process that the target device identifier is consistent with any current device identifier is represented by the following formula:

[0147] (23)

[0148] In formula (23), represents the final judgment result of the comparison operation, represents the number of device identifiers in the list structure, represents the target device identifier contained in the intention vector, represents the i-th device identifier in the device identifier set, represents the i-th device identifier in the device identifier set, represents the identification consistency judgment operation, represents the logical or operation. The control logic of formula (23) is to check whether any identifier in the device list is the same as the target device. If there is one, the result is consistent, otherwise, it is inconsistent.

[0149] When obtaining the target device identifier in the operation intention vector and the device identifier set in the list structure, first, the intention vector needs to be extracted from the user's interactive input. This vector is usually an embedded representation generated by a natural language processing model, such as converting the user query "control the living room light" into a vector form through the BERT (Bidirectional Encoder Representations from Transformers) model, where the target device identifier may be "Light_001". At the same time, the device identifier set in the list structure is pulled from the previously updated device list, such as an array containing multiple identifiers ["Light_001", "AC_002", "Sensor_003"].

[0150] The device identifier sequence acquisition unit is used to group and sort the current device identifiers in the list structure based on dynamic parameters through a clustering algorithm if the preliminary matching result shows that the target device identifier is inconsistent with the current device identifier in the list structure, to obtain a sorted device identifier sequence.

[0151] The preliminary matching result between the target device identifier and the current device identifier is obtained using the following formula:

[0152] (twenty four)

[0153] In formula (24), This indicates the preliminary matching result between the target device identifier and the current device identifier. Indicates the target device identifier. This represents the current device identifier in the list structure. The threshold value represents the matching threshold. When the matching result is 0, it indicates that the two device identifiers are inconsistent. The control logic of formula (24) is to calculate the difference between the two device identifiers. If the difference is within the threshold, it is considered a match; otherwise, it is not a match.

[0154] The sorted device identifier sequence is obtained using the following formula:

[0155] (25)

[0156] In formula (25), This represents the sorted sequence of device identifiers. Indicates the sorted order of the first... The device identifier of the bit. Indicates the total number of device identifiers. This represents a sorting and permutation function. This represents a scoring function based on dynamic parameters. This represents the set of dynamic parameters. The control logic of formula (25) is to first score each device according to its dynamic parameters, and then rank the devices according to their scores. The permutation function is responsible for determining the device corresponding to each ranking.

[0157] The process of determining whether a target device identifier matches any current device identifier in the list structure using a comparison operation is specifically implemented through string matching or hash table lookup. For example, Python's set data structure can be used to quickly check if "Light_001" is in the set. If it exists, the initial match result is considered a match; otherwise, it is considered a mismatch. This allows for rapid response to user intent and avoids delays caused by invalid operations. For instance, if the initial match result indicates that the target device identifier does not match the current device identifier in the list structure—for example, "AC_003" in the user intent vector is not in the list ["AC_001", "AC_002"]—this may be due to a newly added device or a changed identifier. In this case, the system will proceed to subsequent processing to ensure an accurate match.

[0158] When grouping and sorting the current device identifiers in the list structure based on dynamic parameters, the clustering algorithm is an unsupervised learning method, such as the K-means algorithm, which groups devices according to dynamic parameters such as temperature, power, and other feature vectors. Specifically, first, collect the dynamic parameter dataset of each identifier, for example, the parameters of "AC_001" [temperature: 24, wind speed: 2] and "AC_002" [temperature: 26, wind speed: 3], then calculate the Euclidean distance to form clusters, such as grouping air conditioners with similar temperatures into a group, and then based on the average value within the cluster or a custom sorting rule such as descending order of power, the sorted device identifier sequence ["AC_002", "AC_001"] is obtained. This helps to prioritize high-load devices and achieve more efficient resource allocation, such as timely adjusting abnormal sensor groups to prevent failures, and can bring optimization effects to device management in business, such as reducing energy consumption.

[0159] A device position index determination unit is configured to use an index mapping operation to compare the target device identifier with the current device identifier in the sorted device identifier sequence one by one, and determine the matching device position index.

[0160] The index value of the first matching position is found by comparing one by one according to the following formula:

[0161] (26)

[0162] In formula (26), represents the device position index that matches successfully, represents the target device identifier to be matched, represents the device identifier at the th position in the sorted device identifier sequence, represents the total length of the device identifier sequence. The control logic of formula (26) is to find from the beginning of the sorted device list, the first position that matches the target device, which is the index to be found.

[0163] When using the index mapping operation to compare the target device identifier with the current device identifier in the sorted device identifier sequence one by one, the index mapping operation is a key-value association mechanism that creates a mapping table from the identifier to the position, such as the device identifier sequence ["Monitor_004", "Sensor_003"] corresponding to the index {"Monitor_004": 0, "Sensor_003": 1}, and then comparing the target "Monitor_004" one by one to determine its position index as 0. This ensures that even if the initial match fails, the potential correspondence can be found through sequence positioning, improving the robustness of the system.

[0164] Furthermore, the 5G AI router integrating voice dialogue and screen display provided in this embodiment includes a synchronous execution command generation module 50 comprising a preliminary synchronization parameter sequence determination unit, a synchronization parameter sequence acquisition unit, and a synchronous execution command sequence generation unit. The preliminary synchronization parameter sequence determination unit is used to obtain a set of device control parameters from the device location index and compare the set of control parameters with the parameter requirements in the voice command one by one using a real-time synchronization mechanism to determine the preliminary synchronization parameter sequence.

[0165] The following formula is used to obtain the weighted average set of device control parameters from the device location index:

[0166] (27)

[0167] In formula (27), Indicates the first A set of control parameters for each device location This indicates the total number of parameters for the device at that location. Indicates the first The weighting coefficients of each parameter. Indicates position First The control parameter values ​​of each device. The control logic of formula (27) is to calculate the weight of each parameter of the device position, add them up and then average them to obtain the control parameter of this position.

[0168] The following formula is used to perform a one-to-one comparison and calculation between the control parameters and the voice command parameters:

[0169] (28)

[0170] In formula (28), Indicates the first Real-time synchronization matching degree of each parameter, Indicates the first in the voice command The required values ​​for each parameter, Indicates the first in the set of control parameters The current values ​​of the parameters, Synchronization factor represents the parameter type. The control logic of formula (28) is "parameter difference quantization + synchronization factor adaptation", the core of which is to calculate the degree of matching between the control parameters and the voice command requirements.

[0171] The following formula is used to determine the rules for generating the initial synchronization parameter sequence:

[0172] (29)

[0173] In formula (29), Indicates the first in the initial synchronization parameter sequence value of the element, denotes the total number of alignment results, denotes the sequence weight of the alignment result, denotes the matching result of the position of the alignment result, denotes the step function, denotes the synchronization threshold. The control logic of formula (29) is to calculate the difference between the parameter requirement and the actual value, normalize it with the maximum value, and then adjust it according to the parameter type to get the matching degree of this parameter.

[0174] The process of obtaining the device control parameter set from the device location index first involves querying the device database based on the previously determined location index. Assuming that the device location index is 2, corresponding to the living room air conditioner device, the system will extract the control parameter set of this device from the cloud storage, including the temperature set value such as 24 degrees Celsius, the wind speed mode such as medium, and the timing switch state such as off. These parameters are presented in a structured array form to ensure that subsequent operations can quickly access specific values, thereby providing basic data support for voice control. For example, when the user issues the voice instruction "set the air conditioner temperature to 22 degrees", the specific implementation of comparing the control parameter set with the parameter requirement in the voice instruction one by one using the real-time synchronization mechanism is completed through an event-driven synchronization framework. Specifically, this mechanism uses message queues such as Kafka to transmit data in real time. First, parse the voice instruction to extract the parameter requirement such as temperature 22 degrees and wind speed automatic. Then, compare the current parameter set of the device such as temperature 25 degrees and wind speed high with it, and match them one by one to generate a preliminary synchronization parameter sequence. For example, the temperature item in the sequence is [current: 25, requirement: 22], and the wind speed item is [current: high, requirement: automatic]. In this way, a preliminary differentiated sequence is formed, which is convenient for further processing.

[0175] The synchronization parameter sequence acquisition unit is used to determine whether the deviation meets the requirements by a preset threshold value if there is a deviation between the parameter values in the preliminary synchronization parameter sequence and the voice instruction requirement, and obtain the adjusted synchronization parameter sequence.

[0176] The deviation percentage between the preliminary synchronization parameter and the voice instruction requirement is calculated by the following formula:

[0177] (30)

[0178] In formula (30), denotes the deviation rate of the parameter, denotes the preliminary synchronization parameter value, denotes the target parameter value of the voice instruction. The control logic of formula (30) is to calculate how much the preliminary parameter deviates from the target value, and then convert the deviation into a percentage based on the target value to reflect the proportion of deviation.

[0179] The generation process of the adjusted synchronization parameter sequence is described by the following formula:

[0180] (31)

[0181] In formula (31), denotes the adjusted synchronization parameter, denotes the original synchronization parameter, denotes the adjustment intensity coefficient, denotes the target parameter value, denotes the deviation value of the parameter, denotes the threshold value, denotes the step function, which adjusts the parameter when the deviation exceeds the threshold value. The control logic of formula (31) is "deviation threshold judgment + conditional parameter adjustment", and the core is to correct the original synchronization parameter based on the target value when the parameter deviation exceeds the threshold value.

[0182] If the parameter value in the preliminary synchronization parameter sequence deviates from the requirement of the voice instruction, the deviation is judged by the preset threshold value to determine whether it meets the requirement. First, the threshold value needs to be defined, such as considering that a temperature deviation of less than 3 degrees is acceptable. Specifically, for a sensor device, if the sequence shows that the current humidity is 60% and the instruction requires 55%, the deviation is 5%, and the preset threshold value is 4%, it is judged as not meeting the requirement. At this time, the system will adjust according to the deviation direction, such as calculating the adjustment value by linear interpolation method to gradually reduce the humidity to 56% as the adjusted value, and finally obtain the adjusted synchronization parameter sequence such as [humidity: 56%, temperature: 23 degrees], which helps to maintain the stable operation of the device.

[0183] The synchronization execution command sequence generation unit is used to convert the parameter sequence into an executable command sequence by using the command generation algorithm based on the adjusted synchronization parameter sequence, and generate a synchronization execution command sequence.

[0184] The following formula is used to represent the conversion of all adjusted parameters into a complete executable command sequence by the command generation algorithm:

[0185] (32)

[0186] In formula (32), denotes the generated synchronization execution command sequence, denotes the total length of the parameter sequence, denotes the adjusted synchronization parameter, denotes the command generation algorithm function. The control logic of formula (32) is "parameter traversal + command conversion integration", and the core is to convert each adjusted parameter into a command one by one, and then integrate it into a complete executable sequence.

[0187] According to the adjusted synchronization parameter sequence, the specific way of converting the parameter sequence into an executable command sequence by using the command generation algorithm is to use a rule-based algorithm engine, specifically, this command generation algorithm first maps the sequence such as [brightness: 50%, alarm: on] to the device API (Application Programming Interface, application programming interface) interface, and then generates JSON format commands such as {“action”:“setBrightness”,“value”:50} and {“action”:“enable alarm”}, thereby forming a synchronous execution command sequence, these commands can be directly issued to the device end for execution, and can bring the optimization effect of real-time response to user demand in business, such as quickly adjusting the monitor parameter in an emergency to improve patient safety.

[0188] Further, the 5G AI router integrating voice dialogue and screen display provided by the embodiment, the network state feedback data acquisition module 60 includes a display content updating unit and a network state feedback data acquisition unit, wherein the display content updating unit is used to update the content displayed on the user interface device according to the synchronous execution command sequence.

[0189] The updated user interface display state is obtained by the following formula:

[0190] (33)

[0191] In formula (33), denotes the updated user interface display state, denotes the command sequence at the current moment, denotes the synchronization execution state, denotes the current user interface state, denotes the interface updating function. The control logic of formula (33) is "multi-state input + interface updating function driven", and the core is to combine the current command, execution state and interface state to generate a new interface display state.

[0192] The process of updating the content displayed on the user interface device based on the synchronously executed command sequence first involves parsing the command sequence into visual data elements. For example, when the command sequence includes adjusting the brightness of the living room lights to medium and the air conditioner temperature to 22 degrees Celsius, the system maps these parameters to icons and numerical displays on the user interface. Specifically, this update is implemented through a front-end framework such as React. Key fields such as brightness and temperature values ​​are extracted from the command sequence and then dynamically rendered into the interface components to ensure that the user can see the changes in real time. For example, the temperature slider on the interface moves from 25 to 22, while the light icon brightens from dark. This completes the synchronous update from command to display.

[0193] The network status feedback data acquisition unit is used to send instructions to the distributed network to adjust the traffic of the target device and obtain the network status feedback data after execution.

[0194] The strength of the traffic adjustment command issued to the distributed network is derived using the following formula:

[0195] (34)

[0196] In formula (34), Indicates at time Adjusting the strength of traffic control commands sent to the distributed network. Indicates the total number of target devices. Indicates the first The weighting coefficient of each device Indicates the first Each device at time The amount of traffic change that needs to be adjusted. This represents the amplification factor for flow adjustment. The adjustment parameter represents the target response. This represents the target flow baseline value that is expected to be achieved. The control logic of formula (34) is "weighted summation of equipment flow + target baseline correction", the core of which is to combine the equipment flow demand and the target baseline to calculate the final flow adjustment command intensity.

[0197] The network status feedback data after execution is obtained using the following formula:

[0198] (35)

[0199] In formula (35), Indicates a delay after the instruction is executed. Network status feedback data obtained over time, This represents the total number of feedback nodes in the network. Indicates the first The priority weight of each feedback node Indicates the first a node at time a network state value at time a gain coefficient representing feedback data, a time delay decay constant, a time delay representing instruction execution to feedback acquisition, a natural constant. The control logic of formula (35) is "multi-node state weighting + time decay correction", and the core is to combine the priority of the feedback node and the time delay to calculate the feedback data of the network state.

[0200] The implementation of issuing instructions to the distributed network to adjust the target device traffic is handled by the API gateway. Specifically, when it is necessary to adjust the network traffic of the in-vehicle entertainment device, the system generates an instruction such as limiting the video stream to 2 Mbps, and then issues it to the distributed nodes such as edge servers. These nodes allocate bandwidth resources in real time according to the instruction to obtain the network state feedback data after execution, for example, the feedback includes the current traffic usage rate such as 80% and the delay time such as 50ms. These data are returned in JSON format for subsequent analysis.

[0201] The embodiment discloses a 5G AI router integrating voice dialogue and screen display. The pure voice signal sequence is obtained by filtering noise through the microphone array beamforming technology, and the neural network model is used to analyze and generate an accurate operation intention vector. The device list structure is updated in real time by pulling dynamic parameters from the distributed network. When the intention vector target device identifier does not match the list position, the clustering algorithm is used to intelligently group and sort the list to quickly determine the matching device position index. Then, the control parameters are extracted to align the voice requirements through the real-time synchronization mechanism, and a synchronous execution command sequence is generated. At the same time, the user interface is updated and the network instruction is issued to adjust the device traffic, and finally the network state feedback data after execution is obtained. The embodiment realizes efficient and accurate matching and control of voice instructions and dynamic devices, improves the robustness and response speed of far-field voice interaction, and ensures seamless execution of device operation in complex network environments.

[0202] Although the preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to these embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be interpreted as including all changes and modifications falling within the scope of the present application. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A 5G AI router integrating voice dialog and screen display, characterized in that, The application relates to a voice command processing method and device. The method comprises the following steps: An instruction content determination module (10) is used for capturing a voice instruction issued by a user from a far-field voice pickup array, obtaining a preliminary voice signal sequence through microphone array signal processing, enhancing directivity pickup by using a beam forming technology to filter noise interference, and determining instruction content containing a keyword; An operation intention vector generation module (20) is used for performing semantic analysis on the voice signal sequence by using a neural network model according to the instruction content, and generating a corresponding operation intention vector, wherein the operation intention vector represents the identification and parameter requirement of a target device; A list structure acquisition module (30) is used for acquiring current device list data on a user interface device, pulling dynamic parameter monitoring information from a distributed network in real time, fusing the dynamic parameter monitoring information into the current device list data to update device states, and obtaining a refreshed list structure; A device position index determination module (40) is used for grouping and sorting the list structure by using a clustering algorithm if the target device identification in the operation intention vector does not match a position in the list structure, and determining a matched device position index; A synchronous execution command generation module (50) is used for extracting a control parameter from the device position index, aligning the control parameter with the requirement of the voice instruction by using a real-time synchronization mechanism, and generating a synchronous execution command sequence; A network state feedback data acquisition module (60) is used for updating the content of the user interface device by using the synchronous execution command sequence, issuing an instruction to a distributed network to adjust the flow of a target device, and obtaining network state feedback data after execution. The operation intention vector generation module (20) comprises: A semantic analysis sequence acquisition unit is used for performing semantic analysis on the voice signal sequence by using a convolutional neural network according to the instruction content, and obtaining a semantic analysis sequence; ; wherein, denotes the final semantic parsing result at the th time step, denotes the set of all possible semantic categories, denotes the total number of convolutional feature maps, denotes the classification weight vector corresponding to the th feature map, denotes the convolutional feature vector at the th time step, denotes the threshold parameter of the th classifier, is an activation function; The semantic analysis sequence is obtained by the following formula: A target device identification sequence acquisition unit is used for obtaining a target device identification from the semantic analysis sequence if the confidence of the semantic analysis sequence exceeds a preset threshold, and obtaining a target device identification sequence; A parameter requirement vector determination unit is used for matching a parameter requirement database established in advance according to the target device identification sequence, and determining a parameter requirement vector; 2. The 5G AI router for integrated voice conversations and screen display of claim 1, wherein, An operation intention vector acquisition unit is used for fusing the target device identification sequence and the parameter requirement vector by using a vector splicing operation if the parameter requirement vector and the target device identification sequence are successfully matched, and obtaining an operation intention vector. The instruction content determination module (10) comprises: A preliminary voice signal sequence acquisition unit is used for obtaining a far-field voice signal sequence captured by a microphone array, wherein the far-field voice signal sequence is a voice instruction issued by a user, and a preliminary voice signal sequence is obtained; A directional voice signal sequence acquisition unit is used for performing delay-sum beam forming processing according to the preliminary voice signal sequence, and judging a directional enhancement signal if the signal energy of the preliminary voice signal sequence exceeds a preset threshold, and obtaining a directional voice signal sequence. The preliminary speech feature sequence acquisition unit is configured to acquire spectral features and time domain features from the directional speech signal sequence to obtain a preliminary speech feature sequence; The instruction content determination unit is configured to determine instruction content by processing the preliminary speech feature sequence using an end-to-end speech recognition model, and determine the instruction content if the preliminary speech feature sequence contains a keyword.

3. The 5G AI router for integrated voice dialogue and screen display of claim 1, wherein, In the target device identifier sequence acquisition unit, the confidence of the semantic analysis sequence is obtained by the following formula: ; wherein, represents a confidence of a semantic parse sequence, represents a length of a parse sequence, represents a probability of the word given all previous words, the confidence of the entire semantic parse sequence is evaluated by averaging the conditional probabilities of all words in the sequence. The threshold value judgment decision condition for defining the confidence is obtained by the following formula: ; wherein, represents a threshold decision result, represents a confidence value of the current semantic parsing sequence, represents a preset confidence threshold, and when the confidence exceeds the threshold, 1 is returned to execute a subsequent extraction operation, otherwise 0 is returned without execution. The target device identifier sequence is obtained by the following formula: ; Wherein, represents the target device identification sequence extracted, containing device identification to , represents the input semantic analysis sequence, represents the pattern matching rule for identifying and extracting device identification, and all target device identifications are extracted from the semantic sequence by the Extract function.

4. The 5G AI router of integrated voice conversations and screen display according to claim 3, wherein, In the parameter requirement vector determination unit, the parameter requirement vector is obtained by the following formula: ; wherein, represents a determined parameter requirement vector value, represents a dimension number of the parameter requirement, represents an importance coefficient of the th parameter, represents a base value of the th parameter, function represents a step function, represents a matching confidence of the th parameter, represents a confidence threshold value.

5. The 5G AI router of integrated voice conversations and screen display of claim 4, wherein, In the operation intention vector acquisition unit, the matching function of the parameter requirement vector and the target device identifier sequence is defined by the following formula: ; wherein, represents a matching function of a parameter requirement vector and a target device identification sequence, represents a parameter requirement vector, represents a target device identification sequence, represents a matching threshold, a match is successful when the Euclidean distance of two vectors is less than or equal to the threshold, otherwise 0 is returned. The operation intention vector is obtained by the following formula: ; wherein, represents an operation intention vector, and represents a vector concatenation operation, represents the i-th element of the device identification sequence, represents the i-th element of the parameter requirement vector, represents the i-th element of the parameter requirement vector, represents the i-th element of the parameter requirement vector, and respectively represent the dimensions of two vectors.

6. The 5G AI router for integrated voice dialogue and screen display of claim 1, wherein, The list structure acquisition module (30) comprises: The initial device list acquisition unit is configured to acquire current device list data in a user interface device, extract current device identifiers and basic attributes of the current device list data using a database query operation, and obtain an initial device list. The dynamic parameter data set acquisition unit is configured to, if the initial device list contains at least one current device identifier, pull dynamic parameter monitoring information corresponding to the current device identifier in real time through a distributed network interface to obtain a dynamic parameter data set. The updated device list acquisition unit is configured to, according to the dynamic parameter data set, match and integrate the dynamic parameters with the current device identifiers in the initial device list using a data fusion algorithm to obtain an updated device list. The list structure acquisition unit is configured to, if the updated device list conforms to a preset list structure format, load the updated device list to a user interface device through an interface refreshing operation to obtain a refreshed list structure.

7. The 5G AI router for integrated voice dialogue and screen display of claim 1, wherein, The device position index determination module (40) comprises: The preliminary matching result acquisition unit is configured to acquire a target device identifier in the operation intention vector and a device identifier set in the list structure, determine whether the target device identifier is consistent with any current device identifier in the list structure using a comparison operation, and obtain a preliminary matching result. The device identifier sequence acquisition unit is configured to, if the preliminary matching result indicates that the target device identifier is not consistent with the current device identifier in the list structure, group and sort the current device identifiers in the list structure based on dynamic parameters using a clustering algorithm to obtain a sorted device identifier sequence. The device position index determination unit is configured to compare the target device identifier with the current device identifiers in the sorted device identifier sequence one by one using an index mapping operation to determine a matched device position index.

8. The 5G AI router for integrated voice dialogue and screen display of claim 1, wherein, The synchronous execution command generation module (50) comprises: The preliminary synchronization parameter sequence determination unit is configured to acquire a device control parameter set from the device position index, compare the control parameter set with parameter requirements in a voice instruction one by one using a real-time synchronization mechanism to determine a preliminary synchronization parameter sequence. The synchronization parameter sequence acquisition unit is configured to, if the parameter value in the preliminary synchronization parameter sequence deviates from the requirement of the voice instruction, determine whether the deviation meets the requirement by using a preset threshold, and obtain an adjusted synchronization parameter sequence. The synchronization execution command sequence generation unit is configured to, according to the adjusted synchronization parameter sequence, convert the parameter sequence into an executable command sequence by using a command generation algorithm, and generate a synchronization execution command sequence.

9. The 5G AI router for integrated voice dialogue and screen display of claim 1, wherein, The network state feedback data acquisition module (60) comprises: The display content update unit is configured to update the content displayed on the user interface device according to the synchronization execution command sequence. The network state feedback data acquisition unit is configured to issue an instruction to the distributed network to adjust the target device flow, and obtain the executed network state feedback data.

Citation Information

Patent Citations

  • Search engine for touch equipment and method

    CN102207960A

  • Voice interaction method and device, and storage medium

    CN119580707A