A 5G communication terminal interactive control system based on AI voice recognition
By employing environmental adaptive preprocessing and edge-cloud collaborative recognition technologies, the purity of voice data and the accuracy of command generation in the AI voice recognition system have been improved. This has solved the problems of insufficient environmental adaptability and network status flexibility in existing technologies, and enabled efficient and reliable voice interaction control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU TAIYUAN ELECTRIC CO LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing AI voice recognition and 5G communication terminal interaction systems have shortcomings in environmental adaptability, network status flexibility, protocol compatibility, and data transmission efficiency and stability. These shortcomings result in low voice recognition accuracy, command transmission delays, and insufficient response data integrity, affecting the real-time performance and reliability of user interaction.
The system performs environmental adaptive preprocessing through a voice acquisition and processing module, combines real-time status information from the communication network for end-to-cloud collaborative recognition, generates and encapsulates control commands that can be recognized by the target device, and improves the system's adaptability and reliability through a multimodal interactive feedback mechanism.
It significantly improves the purity of voice data and the accuracy of command generation, ensures the compatibility and transmission efficiency of data packets, and enhances the smoothness of user interaction and the stability of the system.
Smart Images

Figure CN122135704A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart terminal technology, and in particular to a 5G communication terminal interactive control system based on AI voice recognition. Background Technology
[0002] In existing technologies for AI speech recognition and interaction with 5G communication terminals, there are significant shortcomings in the environmental adaptability preprocessing capabilities of speech data. It is difficult to dynamically capture dynamic noise characteristics in complex scenarios, and the suppression effect on environmental background noise and non-human voice interference components in the original speech data stream is limited. This results in a lot of interference information remaining in the purified speech data, which directly affects the accuracy of subsequent speech recognition, leading to a high error rate in text information conversion and creating potential problems for subsequent command generation and device interaction.
[0003] Existing technologies fail to achieve end-to-cloud collaborative identification and scheduling based on the real-time status of communication networks. Identification path selection lacks flexibility, and the identification mode cannot be reasonably matched to fluctuations in network transmission latency and signal strength, resulting in a trade-off between identification efficiency and quality. Furthermore, the instruction conversion process suffers from insufficient adaptability to different target device communication protocols, poor data packet encapsulation compatibility and accuracy, and a lack of scientific priority determination and channel optimization mechanisms for data transmission. This leads to instruction transmission delays and inadequate verification of response data integrity, ultimately resulting in insufficient real-time performance and reliability of user-device interaction. The overall operational efficiency and stability of the interactive control system need improvement. Therefore, improving the efficiency of communication terminal interaction has become an urgent problem to be solved. Summary of the Invention
[0004] To achieve the above objectives, the present invention provides a 5G communication terminal interactive control system based on AI voice recognition, characterized in that the system includes a voice acquisition and processing module, a voice format conversion module, an instruction generation module, an instruction conversion module, a data transmission module, and an interactive feedback module, wherein: The voice acquisition and processing module is used to receive the user's raw voice data stream through a terminal of a communication network, and to perform environmental adaptive preprocessing on the raw voice data stream to obtain the user's cleaned voice data. The voice format conversion module is used to perform end-to-cloud collaborative recognition on the purified voice data based on the real-time status information of the communication network to obtain the user's text information; The instruction generation module is used to perform semantic recognition on the text information based on the context of the purified voice data, obtain the user's voice intent, and encode the voice intent into the initial control instruction of the communication network. The instruction conversion module is used to encapsulate the initial control instruction based on the communication protocol characteristics of the target device to obtain a data packet that can be identified by the target device. The data transmission module is used to send the identifiable data packet to the target device through the communication network and to monitor the response data of the target device in real time. The interactive feedback module is used to perform multimodal interactive processing on the response data to obtain the user's interactive data.
[0005] In a preferred embodiment, when the voice acquisition and processing module receives the user's raw voice data stream through a terminal on a communication network and performs environment-adaptive preprocessing on the raw voice data stream to obtain the user's cleaned voice data, it is specifically used for: The user's raw voice data stream is received through a microphone array set in the terminal of the communication network. The ambient background sound is separated from the original voice data stream to obtain the user's ambient audio segment; Real-time spectrum analysis of the environmental audio segment is performed to obtain the dynamic noise characteristics of the user; Based on the dynamic noise characteristics, the original voice data stream is filtered to obtain the user's first-level purified voice data; Spatial filtering is performed on the first-level purified voice data to obtain the user's second-level purified voice data. The user's purified voice data is obtained by suppressing non-human voice interference components in the second-level purified voice data.
[0006] In a preferred embodiment, when the voice format conversion module performs end-to-cloud collaborative recognition on the purified voice data based on the real-time status information of the communication network to obtain the user's text information, it is specifically used for: The transmission delay, signal strength, and available bandwidth of the communication network are used as the real-time status information of the communication network. Based on the transmission delay and the signal strength, the stability of the communication network is evaluated to obtain a network stability rating for the communication network. Based on the real-time status information, the feasibility of transmitting the purified voice data is evaluated, and a feasibility rating for transmitting the purified voice data is obtained. The network stability rating and the transmission feasibility rating are encapsulated and encoded to obtain the user's transmission path decision instruction; Based on the transmission path decision instruction, the purified voice data is format-converted to obtain the user's text information.
[0007] In a preferred embodiment, when the voice format conversion module executes the transmission path decision instruction to convert the format of the purified voice data to obtain the user's text information, it is specifically used for: When the transmission path decision instruction indicates local recognition, the purified voice data is locally decoded to obtain the user's local recognition text, and the local recognition text is used as the user's text information; When the transmission path decision instruction indicates cloud recognition, the upload strategy for the purified voice data is determined based on the encoding format of the purified voice data and the real-time status information. According to the upload strategy, and through the communication network, the purified voice data is uploaded to the cloud-based voice recognition service of the communication network to obtain the user's cloud-based recognized text, and the cloud-based recognized text is used as the user's text information; When the path decision instruction indicates collaborative recognition, structural analysis is performed on the purified speech data to obtain the key speech segments and auxiliary speech segments of the purified speech data; Speech recognition is performed on the key speech segments to obtain the key text of the purified speech data; Based on the real-time status information, the auxiliary speech segment is uploaded to the cloud speech recognition service to obtain the auxiliary text of the purified speech data; The key text and the auxiliary text are fused and verified to obtain the user's text information.
[0008] In a preferred embodiment, when the instruction generation module performs semantic recognition on the text information based on the context of the purified voice data to obtain the user's voice intent, and encodes the voice intent into initial control instructions for the communication network, it is specifically used for: Lexical analysis is performed on the text information to obtain the control action words and target device words of the text information; Based on the context of the purified voice data, the referential relationship of the target device words is parsed to obtain the target object of the text information; Based on the control action words and the target object, the user's control purpose at the current moment is inferred, and the user's voice intent is obtained; Based on the standardized instruction format of the target device, the voice intent is converted into structured instruction elements of the target device; Based on the data transmission format specified by the communication network, the structured instruction elements are arranged and encapsulated to obtain the initial control instructions of the communication network.
[0009] In a preferred embodiment, when the instruction conversion module performs data encapsulation on the initial control instruction based on the communication protocol characteristics of the target device to obtain a data packet recognizable by the target device, it specifically performs the following: Parse the device identifier of the target device, query the communication protocol description file bound to the device identifier, and extract features from the communication protocol description file to obtain the protocol feature set of the target device; Based on the instruction structure specifications in the protocol feature set, the structured instruction elements in the initial control instruction are serialized and rearranged to obtain the intermediate instruction data of the initial control instruction. Based on the data field encoding rules in the protocol feature set, the instruction parameters in the intermediate instruction data are serialized and encoded to obtain the data encoding format of the target device. The data encoding format is validated to obtain the optimized data encoding format for the target device; The optimized data encoding format and the intermediate instruction data are encapsulated into a data packet that is recognizable by the target device.
[0010] In a preferred embodiment, when the instruction conversion module executes the instruction structure specification based on the protocol feature set to serialize and rearrange the structured instruction elements in the initial control instruction to obtain the intermediate instruction data of the initial control instruction, it is specifically used for: The initial control command is parsed in a structured manner to obtain the structured instruction elements of the initial control command; Based on the instruction structure specification, the structured instruction elements in the initial control instruction are length aligned to obtain the optimized instruction elements of the initial control instruction. Based on the instruction structure specification, the optimized instruction elements are syntactically assembled to obtain the intermediate instruction data of the initial control instruction.
[0011] In a preferred embodiment, when the data transmission module sends the identifiable data packet to the target device via the communication network and monitors the response data of the target device in real time, it is specifically used for: The initial control command is prioritized to determine its urgency level. Based on the urgency level and the current load status of the communication network, determine the transmission priority and transmission channel type of the identifiable data packet; Based on the transmission channel type, the identifiable data packet is encapsulated at the link layer to obtain the network transmission data of the communication network; The network transmission data is sent to the target device, and the initial response data of the target device is monitored at the response port of the target device. The initial response data is subjected to integrity verification to obtain the response data of the target device.
[0012] In a preferred embodiment, when the data transmission module performs the step of determining the transmission priority of the identifiable data packet based on the urgency level and the current load state of the communication network, it includes: The urgency level and the current load status of the communication network are normalized to obtain the standard urgency level and standard load values of the communication network. Statistically analyze the historical response success rate of the target device; Based on the standard urgency, the standard load value, and the historical response success rate, the transmission priority of the identifiable data packet is calculated, wherein the formula for calculating the transmission priority is: ; in, Indicates the transmission priority, Indicates the standard urgency. This represents the standard load value. This indicates the historical response success rate. This represents the preset decision curvature adjustment constant.
[0013] In a preferred embodiment, when the interactive feedback module performs multimodal interactive processing on the response data to obtain the user's interactive data, it is specifically used for: The response data is semantically parsed to obtain the instruction execution status and instruction semantic description of the target device, and the feedback type of the target device is determined based on the instruction semantic description; Speech synthesis is performed on the feedback type and the instruction semantic description to obtain the speech broadcast data stream of the target device; Visual rendering is performed on the instruction execution state and the instruction semantic description to obtain visual demonstration data of the target device; By integrating the voice broadcast data stream and the visual demonstration data, the user's interactive data is obtained.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention performs multi-stage purification processing on the original voice data stream through environmental adaptive preprocessing, effectively removing environmental noise and non-human voice interference components, and significantly improving the purity of the voice data; at the same time, it relies on the real-time status information of the communication network to achieve end-to-cloud collaborative recognition, accurately match the voice format conversion path, and combine contextual semantic parsing and intent inference to ensure the accuracy of control command generation, greatly improving the reliability of voice interaction and the effectiveness of command conversion.
[0015] 2. This invention encapsulates and optimizes the format of instructions by utilizing the communication protocol characteristics of the target device, ensuring the compatibility and identifiability of data packets; it improves the efficiency and stability of data transmission by prioritizing and dynamically selecting transmission channels, in conjunction with response data integrity verification; and it integrates voice broadcasting and visual demonstration data through a multimodal interactive feedback mechanism, allowing users to clearly perceive the instruction execution status, further enhancing the system's adaptability and the smoothness of the user interaction experience. Attached Figure Description
[0016] Figure 1 A system architecture diagram of a 5G communication terminal interactive control system based on AI voice recognition is provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments belong to some, but not all, embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “said” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.
[0019] Depending on the context, the word "if" or "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0020] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0021] In practice, the server-side equipment deployed in a 5G communication terminal interaction control system based on AI voice recognition may consist of one or more devices. This AI-based 5G communication terminal interaction control system can be implemented as a service instance, a virtual machine, or hardware devices. For example, it can be implemented as a service instance deployed on one or more devices in a cloud node. Simply put, it can be understood as software deployed on a cloud node, providing a 5G communication terminal interaction control system based on AI voice recognition for each user terminal. Alternatively, it can be implemented as a virtual machine deployed on one or more devices in a cloud node, with application software installed to manage each user terminal. Or, it can also be implemented as a server composed of numerous identical or different types of hardware devices, with one or more hardware devices configured to provide a 5G communication terminal interaction control system based on AI voice recognition for each user terminal.
[0022] In terms of implementation, the AI-based voice recognition 5G communication terminal interactive control system and the user terminal are mutually compatible. Specifically, if the AI-based voice recognition 5G communication terminal interactive control system is implemented as an application installed on a cloud service platform, the user terminal acts as a client establishing a communication connection with that application; or if the AI-based voice recognition 5G communication terminal interactive control system is implemented as a website, the user terminal acts as a webpage; or if the AI-based voice recognition 5G communication terminal interactive control system is implemented as a cloud service platform, the user terminal acts as a mini-program within an instant messaging application.
[0023] like Figure 1 The diagram shown is a system architecture diagram of a 5G communication terminal interactive control system based on AI voice recognition, provided by an embodiment of the present invention.
[0024] The AI-based voice recognition-based 5G communication terminal interactive control system 100 described in this invention can be located on a cloud server. In terms of implementation, it can function as one or more service devices, or as an application installed on the cloud (e.g., a mobile service operator's server, server cluster, etc.), or it can be developed into a website. Depending on the functions implemented, the AI-based voice recognition-based 5G communication terminal interactive control system 100 may include a voice acquisition and processing module 101, a voice format conversion module 102, an instruction generation module 103, an instruction conversion module 104, a data transmission module 105, and an interactive feedback module 106. The module described in this invention can also be called a unit, referring to a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.
[0025] In this embodiment of the invention, in a 5G communication terminal interactive control system based on AI voice recognition, each of the above-mentioned modules can be implemented independently and can call other modules. Here, "calling" can be understood as one module connecting to multiple modules of another type and providing corresponding services to those connected modules. In the 5G communication terminal interactive control system based on AI voice recognition provided by this embodiment of the invention, without modifying the program code, the applicability of the architecture of a 5G communication terminal interactive control system based on AI voice recognition can be adjusted by adding modules and directly calling them, achieving cluster-based horizontal expansion to quickly and flexibly expand the 5G communication terminal interactive control system based on AI voice recognition. In practical applications, the above-mentioned modules can be set in the same device or different devices, or they can be set in virtual devices, such as service instances in a cloud server.
[0026] The following describes, with reference to specific embodiments, each component and its specific workflow of a 5G communication terminal interactive control system based on AI voice recognition: The voice acquisition and processing module 101 is used to receive the user's raw voice data stream through a terminal of a communication network, and to perform environmental adaptive preprocessing on the raw voice data stream to obtain the user's cleaned voice data. In this embodiment of the invention, when the voice acquisition and processing module receives the user's raw voice data stream through a terminal on a communication network and performs environment-adaptive preprocessing on the raw voice data stream to obtain the user's cleaned voice data, it is specifically used for: The user's raw voice data stream is received through a microphone array set in the terminal of the communication network. The ambient background sound is separated from the original voice data stream to obtain the user's ambient audio segment; Real-time spectrum analysis of the environmental audio segment is performed to obtain the dynamic noise characteristics of the user; Based on the dynamic noise characteristics, the original voice data stream is filtered to obtain the user's first-level purified voice data; Spatial filtering is performed on the first-level purified voice data to obtain the user's second-level purified voice data. The user's purified voice data is obtained by suppressing non-human voice interference components in the second-level purified voice data.
[0027] The microphone array set up in the terminal of the communication network consists of multiple evenly distributed microphone units. All microphone units work synchronously to capture audio signals in the surrounding space. The voice signal emitted by the user is received by each microphone unit along with other sound signals in the environment. These microphone units transmit the audio signals they have collected to the data aggregation unit. After the signals are synchronized and integrated, a continuous and unprocessed raw voice data stream is formed.
[0028] Based on the fundamental differences between speech signals and ambient background sounds in terms of frequency distribution and signal stability, the original speech data stream is processed into frames of fixed duration. Each frame corresponds to a short period of audio information. By analyzing the duration and frequency fluctuation amplitude of the signals in each frame, signals that are persistent, have smooth frequency changes, and do not change with the user's vocalization are selected. These selected signals are then integrated to form an ambient audio segment.
[0029] The environmental audio segment is divided into multiple consecutive analysis segments according to time sequence. After each analysis segment is input into the spectrum analysis device, the device will perform frequency decomposition on the audio segment to identify all the frequency components contained therein and the signal strength corresponding to each frequency component. The frequency composition changes of the environmental audio and the intensity fluctuations of each frequency component in different time periods are continuously tracked and recorded. These real-time frequency and intensity related information are systematically organized to form a dynamic noise characteristic that can comprehensively reflect the changing law of environmental noise.
[0030] Based on the noise frequency range and intensity parameters determined from the dynamic noise characteristics, the working threshold of the filtering device is adjusted so that the filtering device can accurately identify noise signals in the original voice data stream that completely match the dynamic noise characteristics. During the filtering process, the core frequency components and amplitude characteristics of the user's voice signal in the original voice data stream are strictly preserved, and only the identified noise signals are completely removed from the original voice data stream. The voice data stream obtained after this processing is the first-level purified voice data.
[0031] By utilizing the spatial distribution information of each microphone unit in the microphone array, the propagation direction and spatial angle of the signal in the first-level purified speech data are analyzed to determine the main source direction of the user's speech signal. The spatial filtering device strengthens the reception and amplifies the signal in this specific spatial direction, while attenuating and suppressing residual interference signals from other spatial directions. Through this spatially selective processing, the residual spatial interference components in the first-level purified speech data are further removed to obtain the second-level purified speech data.
[0032] The standard frequency range and amplitude variation interval of the human voice signal are preset. The second-level purified voice data is compared with the preset human voice feature standard segment by segment. Signal segments that exceed the range of the human voice feature standard are filtered out. These segments are non-human voice interference components. The signal intensity of these non-human voice interference components is reduced by a special signal suppression circuit so that they will not affect the subsequent speech recognition processing. After this suppression processing, the pure purified voice data is finally obtained.
[0033] The beneficial effects are that, through a step-by-step, targeted processing procedure, environmental background noise, dynamic noise, spatial interference, and non-human voice interference components in the original voice data stream are removed sequentially, successfully obtaining high-purity purified voice data. This effectively eliminates the adverse effects of various interference factors on subsequent voice format conversion, semantic recognition, and command generation, ensuring the accuracy and effectiveness of data processing in subsequent modules. It provides reliable voice data support for the stable and efficient operation of the entire AI-based 5G communication terminal interactive control system.
[0034] The voice format conversion module 102 is used to perform end-to-cloud collaborative recognition on the purified voice data based on the real-time status information of the communication network to obtain the user's text information; In this embodiment of the invention, when the voice format conversion module performs end-to-cloud collaborative recognition on the purified voice data based on the real-time status information of the communication network to obtain the user's text information, it is specifically used for: The transmission delay, signal strength, and available bandwidth of the communication network are used as the real-time status information of the communication network. Based on the transmission delay and the signal strength, the stability of the communication network is evaluated to obtain a network stability rating for the communication network. Based on the real-time status information, the feasibility of transmitting the purified voice data is evaluated, and a feasibility rating for transmitting the purified voice data is obtained. The network stability rating and the transmission feasibility rating are encapsulated and encoded to obtain the user's transmission path decision instruction; Based on the transmission path decision instruction, the purified voice data is format-converted to obtain the user's text information.
[0035] When the voice format conversion module executes the transmission path decision instruction to convert the format of the purified voice data and obtain the user's text information, it is specifically used for: When the transmission path decision instruction indicates local recognition, the purified voice data is locally decoded to obtain the user's local recognition text, and the local recognition text is used as the user's text information; When the transmission path decision instruction indicates cloud recognition, the upload strategy for the purified voice data is determined based on the encoding format of the purified voice data and the real-time status information. According to the upload strategy, and through the communication network, the purified voice data is uploaded to the cloud-based voice recognition service of the communication network to obtain the user's cloud-based recognized text, and the cloud-based recognized text is used as the user's text information; When the path decision instruction indicates collaborative recognition, structural analysis is performed on the purified speech data to obtain the key speech segments and auxiliary speech segments of the purified speech data; Speech recognition is performed on the key speech segments to obtain the key text of the purified speech data; Based on the real-time status information, the auxiliary speech segment is uploaded to the cloud speech recognition service to obtain the auxiliary text of the purified speech data; The key text and the auxiliary text are fused and verified to obtain the user's text information.
[0036] The communication network status monitoring unit captures the time difference during transmission in real time, which is the transmission delay. At the same time, the signal strength is obtained by continuously detecting the strength of the network signal through the signal sensing component built into the terminal. The available bandwidth is determined by statistically analyzing the total amount of data that the network can successfully transmit per unit time. These three parameters reflecting the current operating status of the network are integrated and used together as the real-time status information of the communication network.
[0037] The system presets a reasonable threshold range for transmission delay and a standard range for signal strength. It compares the real-time detected transmission delay with the preset threshold and determines whether the signal strength is within the standard range. Based on the comparison results and the range in which the signal strength is located, it classifies the network into three levels: stable, relatively stable, and unstable. The network stability rating of the communication network is obtained based on the rating result.
[0038] By combining the transmission latency, signal strength, and available bandwidth in the real-time status information, the network conditions required for purified voice data transmission are analyzed to determine whether the current available bandwidth can meet the transmission volume requirements of purified voice data, whether the transmission latency is within the acceptable range of data transmission, and whether the signal strength is sufficient to support uninterrupted data transmission. Based on these judgment results, three levels are divided into high feasibility, medium feasibility, and low feasibility, and finally, a feasibility rating for the transmission of purified voice data is obtained.
[0039] The network stability rating and transmission feasibility rating are converted into corresponding binary codes. These codes are then combined in a fixed order of "network stability rating code + transmission feasibility rating code". A start identifier and an end identifier are added to form a complete encoding sequence. This encoding sequence is the user's transmission path decision instruction, which is used to clarify the execution path of subsequent speech recognition.
[0040] When the encoded sequence in the transmission path decision instruction corresponds to the local identification identifier, the voice decoding program stored locally on the terminal is started. This program parses the audio signal frame by frame according to the encoding structure of the purified voice data, converts the voice content corresponding to the audio signal into text symbols, and integrates the text after all frames are parsed to obtain the user's local identification text. This local identification text is then directly used as the user's text information.
[0041] When the encoded sequence in the transmission path decision instruction corresponds to the cloud identification identifier, the encoding format of the purified voice data is first determined by the format recognition unit of the terminal. Then, combined with the available bandwidth and transmission delay in the real-time status information, if the available bandwidth is sufficient and the transmission delay is short, a one-time complete upload mode is adopted. If the available bandwidth is limited or the transmission delay is long, a segmented upload mode based on fixed data segments is adopted. The core parameters such as upload mode and upload rate are clearly defined to form the upload strategy for purified voice data.
[0042] According to the established upload strategy, a dedicated data transmission link between the terminal and the cloud-based voice recognition service is established through the 5G communication network. The purified voice data is gradually transmitted to the cloud according to the upload mode. After receiving the data, the cloud-based voice recognition service performs accurate analysis and speech-to-text processing on the audio signal to generate the corresponding text content. After grammar proofreading and semantic verification in the cloud, the user's cloud-recognized text is obtained and used as the user's text information.
[0043] When the encoded sequence in the transmission path decision instruction corresponds to the collaborative identification identifier, the speech structure analysis program is invoked to parse the purified speech data segment by segment, identify the speech segments containing the user's core needs and key control instructions, and designate them as key speech segments. At the same time, speech segments that play a role in supplementing context and providing auxiliary explanations are designated as auxiliary speech segments, thereby obtaining the key speech segments and auxiliary speech segments of the purified speech data.
[0044] The terminal's local speech recognition engine is activated, and key speech segments are input into the engine. The engine converts the audio signals of the key speech segments into text sentence by sentence according to the preset speech-to-text correspondence library. During the conversion process, the accuracy of the core information is ensured. After the conversion is completed, the key text of the purified speech data is generated.
[0045] Referring to the network stability rating and transmission feasibility rating in the real-time status information, when the network status is relatively stable, the auxiliary voice segment is uploaded to the cloud speech recognition service through the communication network. The cloud service performs speech-to-text processing on the auxiliary voice segment to ensure the complete conversion of auxiliary information and obtain auxiliary text with purified voice data.
[0046] The key text and auxiliary text are arranged in chronological order in the original cleaned speech data. The semantic logic of the key text and auxiliary text is checked to ensure that they are consistent. Contextual information related to the key text is added to the auxiliary text. Any recognition biases between the two are corrected to ensure that the integrated text can completely and accurately reflect the user's speech content, and finally obtain the user's text information.
[0047] The beneficial effects include the ability to flexibly switch between local, cloud, and collaborative recognition paths by accurately collecting real-time status information of the communication network and conducting multi-dimensional evaluation. This ensures efficient conversion of purified voice data into text information under different network conditions. Local recognition guarantees rapid response when the network is poor, cloud recognition improves the recognition accuracy of complex voice data, and collaborative recognition balances response speed and recognition quality. At the same time, through multi-stage verification and integration, the accuracy and completeness of text information are effectively improved, providing high-quality data support for subsequent semantic recognition and command generation, and enhancing the adaptability and reliability of the entire interactive control system.
[0048] The instruction generation module 103 is used to perform semantic recognition on the text information based on the context of the purified voice data, obtain the user's voice intent, and encode the voice intent into the initial control instruction of the communication network. In this embodiment of the invention, when the instruction generation module performs semantic recognition on the text information based on the context of the purified voice data to obtain the user's voice intent, and encodes the voice intent into an initial control instruction for the communication network, it is specifically used for: Lexical analysis is performed on the text information to obtain the control action words and target device words of the text information; Based on the context of the purified voice data, the referential relationship of the target device words is parsed to obtain the target object of the text information; Based on the control action words and the target object, the user's control purpose at the current moment is inferred, and the user's voice intent is obtained; Based on the standardized instruction format of the target device, the voice intent is converted into structured instruction elements of the target device; Based on the data transmission format specified by the communication network, the structured instruction elements are arranged and encapsulated to obtain the initial control instructions of the communication network.
[0049] The text information is broken down sentence by sentence, and the text is divided into independent lexical units according to the natural boundaries of semantic expression. Combined with a pre-set lexical classification database, the semantic attributes of each lexical unit are determined, and words that represent specific operational behaviors are selected as control action words. These words directly correspond to the operation instructions of the device. At the same time, words that refer to the specific controlled device are selected as target device words. These words clarify the specific object to which the operation is directed. Finally, lexical parsing is completed to obtain the control action words and target device words of the text information.
[0050] The context of purified voice data includes related voice segments, previous user interaction records, and current usage scenario information. Based on this context, it is analyzed whether the target device word has a referential expression. If the target device word is a referential word such as "it" or "the device", the device name explicitly mentioned in the context is traced back, or it is matched with the unique device type in the scenario. If the target device word is a specific name, its corresponding actual device is directly identified. In this way, the referential relationship is resolved, and the target object of the text information is accurately obtained.
[0051] By associating and matching the specific operation represented by the control action word with the controlled device specified by the target object, the correspondence between the two is clarified. For example, when the control action word is "adjust" and the target object is "brightness of the living room light", the core need of the user is to adjust the brightness of the living room light by semantic association. By comprehensively sorting out the combination logic of this operation and object, the user's control purpose at the current moment can be accurately inferred, and thus the user's voice intent can be obtained.
[0052] Standardized instruction formats corresponding to various types of target devices are pre-stored. These formats clearly define the core fields that the instructions must include, such as operation type, device identifier, and parameter information. Based on the specific content of the voice intent, it is broken down into specific information corresponding to each field in the standardized instruction format. For example, if the voice intent is "turn on the bedroom air conditioner and set the temperature to 25 degrees Celsius", it is broken down into operation type "turn on", device identifier "bedroom air conditioner", and parameter information "temperature 25 degrees Celsius". These decomposed structured information are the structured instruction elements of the target device.
[0053] The data transmission format specified by the communication network defines the order of data fields, syntax rules, and encapsulation requirements. According to the format requirements, the structured instruction elements are sequentially mapped to the respective field positions of the transmission format. Each element is then processed for format adaptation to ensure that each element conforms to the syntax specifications of the transmission format. Subsequently, all structured instruction elements are integrated into a complete data unit according to the specified encapsulation method. This data unit is the initial control instruction of the communication network.
[0054] The beneficial effects are that, through a progressive parsing and conversion process, the core operations and controlled objects in the text information are accurately extracted, effectively solving problems such as semantic ambiguity and unclear referentiality, ensuring accurate recognition of voice intent. At the same time, the instructions are structured and encapsulated according to a standardized format, so that the initial control instructions not only meet the recognition requirements of the target device, but also adapt to the transmission specifications of the communication network. This provides a precise and standardized data foundation for subsequent instruction conversion and data transmission, significantly improving the semantic recognition accuracy and instruction generation reliability of the entire interactive control system.
[0055] The instruction conversion module 104 is used to encapsulate the initial control instruction based on the communication protocol characteristics of the target device to obtain a data packet that can be identified by the target device. In this embodiment of the invention, when the instruction conversion module performs data encapsulation on the initial control instruction based on the communication protocol characteristics of the target device to obtain a data packet identifiable by the target device, it is specifically used for: Parse the device identifier of the target device, query the communication protocol description file bound to the device identifier, and extract features from the communication protocol description file to obtain the protocol feature set of the target device; Based on the instruction structure specifications in the protocol feature set, the structured instruction elements in the initial control instruction are serialized and rearranged to obtain the intermediate instruction data of the initial control instruction. Based on the data field encoding rules in the protocol feature set, the instruction parameters in the intermediate instruction data are serialized and encoded to obtain the data encoding format of the target device. The data encoding format is validated to obtain the optimized data encoding format for the target device; The optimized data encoding format and the intermediate instruction data are encapsulated into a data packet that is recognizable by the target device.
[0056] When the instruction conversion module executes the instruction structure specification based on the protocol feature set and serializes and rearranges the structured instruction elements in the initial control instruction to obtain the intermediate instruction data of the initial control instruction, it is specifically used for: The initial control command is parsed in a structured manner to obtain the structured instruction elements of the initial control command; Based on the instruction structure specification, the structured instruction elements in the initial control instruction are length aligned to obtain the optimized instruction elements of the initial control instruction. Based on the instruction structure specification, the optimized instruction elements are syntactically assembled to obtain the intermediate instruction data of the initial control instruction.
[0057] The terminal's built-in device identifier resolution program reads the target device's unique device identifier, which contains core information such as device type and manufacturer. Based on this identifier, a precise query is performed in the system's device protocol database to locate the communication protocol description file uniquely bound to the device identifier. This file records the target device's communication rules in detail. Subsequently, key information such as instruction structure requirements, data field formats, and encoding standards in the file are extracted and organized to form a protocol feature set covering the core communication characteristics of the target device.
[0058] The startup command parsing tool breaks down the initial control command layer by layer. According to the composition logic of the initial control command, it separates the structured command elements such as operation type, device identifier, and parameter information one by one. Each element corresponds to a clear semantic and function. By classifying and organizing, it ensures that all structured command elements are completely extracted, and finally obtains the structured command elements of the initial control command.
[0059] Referring to the instruction structure specifications in the protocol feature set, the standard length requirement corresponding to each structured instruction element is defined. The actual length of each structured instruction element is compared with the standard length. For elements that are not long enough, placeholders conforming to the specifications are added to the end to meet the length requirement. For elements that exceed the standard length, redundant and invalid information is trimmed to ensure that the length of each structured instruction element strictly conforms to the instruction structure specifications. After adjustment, the optimized instruction elements of the initial control instructions are obtained.
[0060] Based on the element arrangement order and syntactic logic specified in the instruction structure specification of the protocol feature set, the optimized instruction elements are arranged in the standard order of "operation type element + device identifier element + parameter information element". At the same time, necessary connectors and separators are added in accordance with the syntactic rules in the specification to form a logically coherent and syntactically correct whole, and finally the intermediate instruction data of the initial control instruction is assembled.
[0061] The encoding rules of the data fields in the protocol feature set are retrieved to clarify the encoding methods corresponding to different types of instruction parameters. For each instruction parameter in the intermediate instruction data, the corresponding encoding rule is selected according to its data type for conversion, and the original value of the parameter is converted into a binary encoding form that the target device can recognize. After all parameters are encoded, they are integrated according to the order of the parameters in the intermediate instruction data to obtain the data encoding format of the target device.
[0062] The format verification tool is activated, and the encoding length, encoding type, and check bit settings of each field in the data encoding format are checked against the encoding format standard specified in the protocol feature set. If the encoding length is found to be incorrect, it is adjusted to the standard length; if the encoding type is incorrect, it is re-encoded according to the correct rules; if the check bit is missing, it is supplemented and generated. After comprehensive verification and correction, the optimized data encoding format of the target device is obtained.
[0063] According to the data packet structure required by the target device's communication protocol, the optimized data encoding format is used as the header information of the data packet to inform the target device of the data encoding method and parsing rules. The intermediate instruction data is used as the core data body of the data packet, containing key information of the user's control instructions. The header information and the data body are combined in a fixed format through a data encapsulation program to form a complete and recognizable data packet that the target device can directly recognize and parse.
[0064] The beneficial effects are that by accurately analyzing the communication protocol characteristics of the target device and performing targeted processing, the initial control commands are perfectly adapted to the communication protocol of the target device. Length alignment, syntax assembly, and serialization encoding ensure the standardization of command data, and format verification further improves the accuracy of data encoding. The resulting recognizable data packets can be quickly identified and parsed by target devices with different communication protocols, which significantly improves the system's compatibility with various target devices, ensures the accuracy and effectiveness of command transmission, and provides a solid guarantee for the smooth transmission of subsequent data and device interaction.
[0065] The data transmission module 105 is used to send the identifiable data packet to the target device through the communication network and to monitor the response data of the target device in real time. In this embodiment of the invention, when the data transmission module sends the identifiable data packet to the target device through the communication network and monitors the response data of the target device in real time, it is specifically used for: The initial control command is prioritized to determine its urgency level. Based on the urgency level and the current load status of the communication network, determine the transmission priority and transmission channel type of the identifiable data packet; Based on the transmission channel type, the identifiable data packet is encapsulated at the link layer to obtain the network transmission data of the communication network; The network transmission data is sent to the target device, and the initial response data of the target device is monitored at the response port of the target device. The initial response data is subjected to integrity verification to obtain the response data of the target device.
[0066] When the data transmission module performs the step of determining the transmission priority of the identifiable data packet based on the urgency level and the current load status of the communication network, it includes: The urgency level and the current load status of the communication network are normalized to obtain the standard urgency level and standard load values of the communication network. Statistically analyze the historical response success rate of the target device; Based on the standard urgency, the standard load value, and the historical response success rate, the transmission priority of the identifiable data packet is calculated, wherein the formula for calculating the transmission priority is: ; in, Indicates the transmission priority, Indicates the standard urgency. This represents the standard load value. This indicates the historical response success rate. This represents the preset decision curvature adjustment constant.
[0067] Referring to the system's preset instruction priority judgment criteria, the urgency attributes of different types of control instructions are clarified. Instructions related to the activation of core functions of the target device and safety protection are defined as high urgency, routine function adjustment instructions are defined as medium urgency, and non-essential operation instructions such as query and status feedback are defined as low urgency. The functional attributes of the initial control instructions are compared with the judgment criteria one by one to finally obtain the urgency level of the initial control instructions.
[0068] The urgency level of the initial control command is converted into standardized data within a fixed numerical range according to preset rules: high urgency corresponds to a value of 1.0, medium urgency corresponds to a value of 0.5, and low urgency corresponds to a value of 0.1, thus completing the normalization process to obtain the standard urgency level. At the same time, the current bandwidth utilization ratio and data transmission queue length of the communication network are collected and converted according to the rule of "0%-30% load corresponds to 0.2, 31%-60% load corresponds to 0.6, and 61%-100% load corresponds to 1.0", thus completing the normalization process to obtain the standard load value.
[0069] Retrieve the historical interaction records of the target device stored in the system, filter out all records of instructions received and processed by the device in the past 30 days, count the number of valid records in which the device successfully returned execution results or had no response exceptions, divide the number of valid records by the total number of instructions sent to the device during this period, and obtain the historical response success rate reflecting the device's past response situation.
[0070] Combining the core influences of standard urgency, standard load value, and historical response success rate, the transmission priority level of identifiable data packets is comprehensively evaluated, with standard urgency as the primary criterion, standard load value as the secondary adjustment criterion, and historical response success rate as the auxiliary correction criterion. Specifically, data packets with a standard urgency of 1.0, a standard load value ≤ 0.6, and a historical response success rate ≥ 80% are classified as high priority; data packets with a standard urgency of 0.5, a standard load value ≤ 0.8, and a historical response success rate ≥ 60% are classified as medium priority; and all other cases are classified as low priority. This process ultimately yields the transmission priority of identifiable data packets.
[0071] Based on the determined urgency level and the current load status of the communication network, match the corresponding transmission channel type. When the urgency level is high and the standard load value is ≤0.6, select a high-speed transmission channel with fast transmission rate and high stability; when the urgency level is medium and the standard load value is ≤0.8, select a conventional transmission channel with strong adaptability and balanced resource usage; when the urgency level is low or the standard load value is >0.8, select an energy-saving transmission channel with low resource consumption and acceptable transmission delay, and clearly define the transmission channel type of the identifiable data packets.
[0072] According to the link layer protocol requirements corresponding to the determined transmission channel type, a frame header information containing the source address, destination address, and channel identifier is added to the header of the identifiable data packet, and a frame tail field for link layer verification is added to the tail of the data packet. At the same time, the control fields specified by the link layer protocol are embedded, and the identifiable data packet is completely embedded into the link layer data structure. After format regularization and field completion, the network transmission data of the communication network is obtained.
[0073] A dedicated data transmission link matching the transmission channel type is established through the 5G communication network. The network transmission data is sent to the target device step by step according to the link transmission specification. During the data transmission process, the target device's preset response port is continuously monitored. Once the target device returns raw data containing instruction reception confirmation and preliminary feedback on execution status, the monitoring is immediately stopped and the data is retained, which is the target device's initial response data.
[0074] By comparing the initial response data with the standard field specifications of the target device's response data, check whether the initial response data contains all necessary fields such as instruction identifier, execution result, and status code. Verify that the content of each field is complete, without missing or incorrect information, and confirm that the data has not been lost, damaged, or tampered with during transmission. After a comprehensive verification and confirmation that there are no errors, the response data of the target device is obtained.
[0075] After determining the urgency level of the initial control command, the obtained urgency level is normalized to obtain the standard urgency level; the current load status of the communication network is normalized to obtain the standard load value; the past response status of the target device is statistically analyzed to obtain the historical response success rate; the decision curvature adjustment normal number is a preset fixed value.
[0076] The standard urgency level is multiplied by the standard load value, and then the product is divided by the historical response success rate. The result of this division is then subjected to a decision curvature adjustment normal power operation. Through this comprehensive calculation process, the transmission priority of identifiable data packets is fully measured. This ensures that the determination of transmission priority can fully combine the urgency of the instruction, the current load status of the communication network, and the historical response performance of the target device, so as to ensure that data transmission can obtain a reasonable priority order according to the actual situation.
[0077] When the result of normalizing the urgency of the initial control command increases, the final transmission priority will increase, provided that the normalized result of the current load state of the communication network and the historical response success rate of the target device remain unchanged.
[0078] When the result of normalizing the current load state of the communication network increases, the final transmission priority will increase, provided that the normalization result of the urgency of the initial control command and the historical response success rate of the target device remain unchanged.
[0079] When the statistical results of the target device's past response increase, the final transmission priority will decrease, provided that the normalized results of the initial control command urgency and the normalized results of the current load status of the communication network remain unchanged.
[0080] When the preset decision curvature adjustment normal value changes, it will change the magnitude of the change in the product of the standard urgency and the standard load value divided by the historical response success rate, thereby changing the drastic change in transmission priority with the above three factors.
[0081] The beneficial effects are that, through scientific priority determination and transmission channel matching, differentiated transmission scheduling of instructions with different levels of urgency is achieved. While ensuring the priority transmission of urgent instructions, network resources are rationally utilized to avoid congestion. Link layer encapsulation ensures that data is adapted to the transmission channel's transmission specifications. Integrity verification effectively filters out damaged or incomplete response data. Real-time monitoring throughout the process ensures smooth connection between instruction transmission and response feedback, significantly improving the efficiency, stability, and accuracy of data transmission, and providing reliable support for the real-time interactive capabilities of the entire interactive control system.
[0082] The interactive feedback module 106 is used to perform multimodal interactive processing on the response data to obtain the user's interactive data.
[0083] In this embodiment of the invention, when the interactive feedback module performs multimodal interactive processing on the response data to obtain the user's interactive data, it is specifically used for: The response data is semantically parsed to obtain the instruction execution status and instruction semantic description of the target device, and the feedback type of the target device is determined based on the instruction semantic description; Speech synthesis is performed on the feedback type and the instruction semantic description to obtain the speech broadcast data stream of the target device; Visual rendering is performed on the instruction execution state and the instruction semantic description to obtain visual demonstration data of the target device; By integrating the voice broadcast data stream and the visual demonstration data, the user's interactive data is obtained.
[0084] The semantic parsing program is invoked to interpret the response data of the target device field by field, extracting the core information that identifies the result of the instruction execution, and clarifying whether the instruction is in a state of successful execution, execution failure, or execution in progress. This result is the instruction execution status of the target device. At the same time, the text content describing the execution process and result details in the response data is extracted to form an instruction semantic description that can clearly explain the instruction execution status. According to the core purpose of the instruction semantic description, if the description is to inform the user of the result after the instruction is completed, the feedback type is determined to be result feedback; if the description is to explain the reason for the instruction execution error, the feedback type is determined to be error feedback; if the description is to indicate the progress of the instruction execution, the feedback type is determined to be progress feedback.
[0085] The corresponding speech synthesis parameters are matched according to the determined feedback type. The result feedback uses clear and steady intonation parameters, the abnormal feedback uses slightly abrupt intonation parameters, and the progress feedback uses uniform and neutral intonation parameters. The text content describing the semantics of the instruction is input into the speech synthesis device. The device converts the text into an audio signal that simulates human voice sentence by sentence according to the matched intonation parameters. The audio signal is then subjected to noise reduction and volume standardization to ensure that the audio is clear and indistinguishable. Finally, the continuous audio signal is organized according to the preset audio encoding format to form the voice broadcast data stream of the target device.
[0086] Based on the execution status of the instruction, select the corresponding visual elements. Successful execution corresponds to green visual elements and a checkmark icon, failure corresponds to red visual elements and an exclamation mark icon, and execution in progress corresponds to blue visual elements and a progress bar icon. Combine the text content describing the semantics of the instruction with the selected visual elements. The text uses a clear and easy-to-read font and an appropriate font size, arranged in a preset layout of "icon + text". Add simple dynamic effects to the visual elements. When execution is successful, the icon flashes briefly, and when execution is in progress, the progress bar is dynamically filled. After visual effect integration and format standardization, the visual demonstration data of the target device is obtained.
[0087] The voice broadcast data stream and visual presentation data are time-aligned to ensure that the content of the voice broadcast and the information of the visual presentation are completely synchronized in time. For example, when the voice broadcasts "The brightness of the lights has been adjusted to 50%", the corresponding text and a dynamic icon of the brightness adjustment progress are displayed in the visual presentation data. The aligned voice broadcast data stream and visual presentation data are encapsulated into a unified interactive data format. This format supports the terminal device to output audio and visual content at the same time, ultimately forming interactive data that users can receive through both auditory and visual channels.
[0088] The beneficial effects are that the core information of instruction execution is accurately obtained through semantic parsing, and targeted speech synthesis and visual rendering are achieved by combining feedback types. Then, the synchronous feedback of speech and vision is achieved through data integration, thus constructing a multimodal interactive feedback mode. This allows users to clearly and accurately know the status of instruction execution through dual senses, avoiding the problems of untimely or unclear information transmission that may occur in a single feedback mode. It significantly improves the intuitiveness and convenience of user interaction with the system, and enhances the user experience and reliability of the entire interactive control system.
[0089] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0090] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A 5G communication terminal interactive control system based on AI voice recognition, characterized in that, The system includes a voice acquisition and processing module, a voice format conversion module, a command generation module, a command conversion module, a data transmission module, and an interactive feedback module, wherein: The voice acquisition and processing module is used to receive the user's raw voice data stream through a terminal of a communication network, and to perform environmental adaptive preprocessing on the raw voice data stream to obtain the user's cleaned voice data. The voice format conversion module is used to perform end-to-cloud collaborative recognition on the purified voice data based on the real-time status information of the communication network to obtain the user's text information; The instruction generation module is used to perform semantic recognition on the text information based on the context of the purified voice data, obtain the user's voice intent, and encode the voice intent into the initial control instruction of the communication network. The instruction conversion module is used to encapsulate the initial control instruction based on the communication protocol characteristics of the target device to obtain a data packet that can be identified by the target device. The data transmission module is used to send the identifiable data packet to the target device through the communication network and to monitor the response data of the target device in real time. The interactive feedback module is used to perform multimodal interactive processing on the response data to obtain the user's interactive data.
2. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the voice acquisition and processing module receives the user's raw voice data stream through a terminal on a communication network and performs environment-adaptive preprocessing on the raw voice data stream to obtain the user's cleaned voice data, it is specifically used for: The user's raw voice data stream is received through a microphone array set in the terminal of the communication network. The ambient background sound is separated from the original voice data stream to obtain the user's ambient audio segment; Real-time spectrum analysis of the environmental audio segment is performed to obtain the dynamic noise characteristics of the user; Based on the dynamic noise characteristics, the original voice data stream is filtered to obtain the user's first-level purified voice data; Spatial filtering is performed on the first-level purified voice data to obtain the user's second-level purified voice data. The user's purified voice data is obtained by suppressing non-human voice interference components in the second-level purified voice data.
3. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the voice format conversion module performs end-to-end collaborative recognition on the purified voice data based on the real-time status information of the communication network to obtain the user's text information, it is specifically used for: The transmission delay, signal strength, and available bandwidth of the communication network are used as the real-time status information of the communication network. Based on the transmission delay and the signal strength, the stability of the communication network is evaluated to obtain a network stability rating for the communication network. Based on the real-time status information, the feasibility of transmitting the purified voice data is evaluated, and a feasibility rating for transmitting the purified voice data is obtained. The network stability rating and the transmission feasibility rating are encapsulated and encoded to obtain the user's transmission path decision instruction; Based on the transmission path decision instruction, the purified voice data is format-converted to obtain the user's text information.
4. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 3, characterized in that, When the voice format conversion module executes the transmission path decision instruction to convert the format of the purified voice data and obtain the user's text information, it is specifically used for: When the transmission path decision instruction indicates local recognition, the purified voice data is locally decoded to obtain the user's local recognition text, and the local recognition text is used as the user's text information; When the transmission path decision instruction indicates cloud recognition, the upload strategy for the purified voice data is determined based on the encoding format of the purified voice data and the real-time status information. According to the upload strategy, and through the communication network, the purified voice data is uploaded to the cloud-based voice recognition service of the communication network to obtain the user's cloud-based recognized text, and the cloud-based recognized text is used as the user's text information; When the path decision instruction indicates collaborative recognition, structural analysis is performed on the purified speech data to obtain the key speech segments and auxiliary speech segments of the purified speech data; Speech recognition is performed on the key speech segments to obtain the key text of the purified speech data; Based on the real-time status information, the auxiliary speech segment is uploaded to the cloud speech recognition service to obtain the auxiliary text of the purified speech data; The key text and the auxiliary text are fused and verified to obtain the user's text information.
5. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the instruction generation module performs semantic recognition on the text information based on the context of the purified voice data to obtain the user's voice intent, and encodes the voice intent into the initial control instruction of the communication network, it is specifically used for: Lexical analysis is performed on the text information to obtain the control action words and target device words of the text information; Based on the context of the purified voice data, the referential relationship of the target device words is parsed to obtain the target object of the text information; Based on the control action words and the target object, the user's control purpose at the current moment is inferred, and the user's voice intent is obtained; Based on the standardized instruction format of the target device, the voice intent is converted into structured instruction elements of the target device; Based on the data transmission format specified by the communication network, the structured instruction elements are arranged and encapsulated to obtain the initial control instructions of the communication network.
6. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the instruction conversion module performs data encapsulation on the initial control instruction based on the communication protocol characteristics of the target device to obtain a data packet recognizable by the target device, it is specifically used for: Parse the device identifier of the target device, query the communication protocol description file bound to the device identifier, and extract features from the communication protocol description file to obtain the protocol feature set of the target device; Based on the instruction structure specifications in the protocol feature set, the structured instruction elements in the initial control instruction are serialized and rearranged to obtain the intermediate instruction data of the initial control instruction. Based on the data field encoding rules in the protocol feature set, the instruction parameters in the intermediate instruction data are serialized and encoded to obtain the data encoding format of the target device. The data encoding format is validated to obtain the optimized data encoding format for the target device; The optimized data encoding format and the intermediate instruction data are encapsulated into a data packet that is recognizable by the target device.
7. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the instruction conversion module executes the instruction structure specification based on the protocol feature set and serializes and rearranges the structured instruction elements in the initial control instruction to obtain the intermediate instruction data of the initial control instruction, it is specifically used for: The initial control command is parsed in a structured manner to obtain the structured instruction elements of the initial control command; Based on the instruction structure specification, the structured instruction elements in the initial control instruction are length aligned to obtain the optimized instruction elements of the initial control instruction. Based on the instruction structure specification, the optimized instruction elements are syntactically assembled to obtain the intermediate instruction data of the initial control instruction.
8. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the data transmission module sends the identifiable data packet to the target device through the communication network and monitors the response data of the target device in real time, it is specifically used for: The initial control command is prioritized to determine its urgency level. Based on the urgency level and the current load status of the communication network, determine the transmission priority and transmission channel type of the identifiable data packet; Based on the transmission channel type, the identifiable data packet is encapsulated at the link layer to obtain the network transmission data of the communication network; The network transmission data is sent to the target device, and the initial response data of the target device is monitored at the response port of the target device. The initial response data is subjected to integrity verification to obtain the response data of the target device.
9. A 5G communication terminal interactive control system based on AI voice recognition as described in claim 8, characterized in that, When the data transmission module performs the step of determining the transmission priority of the identifiable data packet based on the urgency level and the current load status of the communication network, it includes: The urgency level and the current load status of the communication network are normalized to obtain the standard urgency level and standard load values of the communication network. Statistically analyze the historical response success rate of the target device; Based on the standard urgency, the standard load value, and the historical response success rate, the transmission priority of the identifiable data packet is calculated, wherein the formula for calculating the transmission priority is: ; in, Indicates the transmission priority, Indicates the standard urgency. This represents the standard load value. This indicates the historical response success rate. This represents the preset decision curvature adjustment constant.
10. The 5G communication terminal interactive control system based on AI voice recognition as described in claim 1, characterized in that, When the interactive feedback module performs multimodal interactive processing on the response data to obtain the user's interactive data, it is specifically used for: The response data is semantically parsed to obtain the instruction execution status and instruction semantic description of the target device, and the feedback type of the target device is determined based on the instruction semantic description; Speech synthesis is performed on the feedback type and the instruction semantic description to obtain the speech broadcast data stream of the target device; Visual rendering is performed on the instruction execution state and the instruction semantic description to obtain visual demonstration data of the target device; By integrating the voice broadcast data stream and the visual demonstration data, the user's interactive data is obtained.