An AI Voice Processing Method and System Based on Intelligent Devices
Through intelligent devices, voice information processing, intent recognition and scenario analysis, combined with AI big model inference and security verification of the Internet of Things platform, the problem of low recognition accuracy of smart devices in diverse usage scenarios and multi-user environments is solved, and efficient and accurate device control and data interaction are achieved.
Patent Information
- Application Number
- CN202510414655.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The speech recognition and natural language understanding technology of existing smart devices has low recognition accuracy in diverse usage scenarios and multi-user environments, resulting in device answers that do not meet users' actual needs.
Through intelligent devices, obtain the voice information uploaded by users, conduct semantic recognition and natural language understanding, obtain user intentions, and combine user identity information and usage scenarios for identity verification and scene extraction, generate operation commands and multimodal feature parameters, upload them to the Internet of Things platform for AI big model inference, generate execution instructions, and establish device connections through security verification to perform operations.
It improves the recognition accuracy of the voice processing system, realizes more accurate and intelligent device control and data interaction, and improves the convenience and intelligence level of user interaction.
Smart Images

Figure CN119920256B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing, and in particular to an AI speech processing method based on intelligent devices. Background Art
[0002] The use of intelligent devices is becoming more and more popular. They are convenient to use and have powerful functions. They can automatically complete various settings according to the conditions set by users, which is very convenient. With the development of language processing technology, more and more intelligent devices are equipped with language recognition functions, such as voice dialing of smart watches. Speech recognition and natural language understanding technology is one of the most important technologies in AI. It enables machines to convert speech signals into corresponding texts or commands through the recognition and understanding process, while understanding the problems and needs of users and giving answers or actions. Simply put, it enables machines to understand people's intentions and needs while allowing people to hear what the machine says, understand what it expresses, and be able to meet people's needs.
[0003] In an existing technology, the language recognition and natural language understanding technology of intelligent devices is based on the recognition and understanding of speech. Although functions such as intelligent question answering and human-machine dialogue can be realized to a certain extent to meet the needs of users. However, the usage scenarios of intelligent devices are becoming more and more diverse, and there are also situations where multiple users use them. The existing technology only recognizes speech, which will cause the answers of intelligent devices not to meet the actual needs of users and the recognition accuracy is low. Summary of the Invention
[0004] The present invention provides an AI speech processing method and system based on intelligent devices to improve the recognition accuracy of the speech processing system.
[0005] In a first aspect, to solve the above technical problems, the present invention provides an AI speech processing method based on intelligent devices, including:
[0006] The intelligent device obtains the voice information uploaded by the user;
[0007] The intelligent device performs semantic recognition and natural language understanding according to the voice information to obtain the user's intention;
[0008] The intelligent device authenticates the user according to the user's intention and the voice information and extracts the user's usage scenario to obtain the user identity information and usage scenario;
[0009] The intelligent device performs text analysis according to the user's intention to obtain a candidate set of operation commands, eliminates the operation commands in the candidate set of operation commands that do not conform to the user identity information and the usage scenario, obtains the target operation command, and extracts multi-modal feature parameters according to the target operation command;
[0010] The intelligent device uploads the operation command and the multi-modal feature parameters to the Internet of Things platform. The Internet of Things platform obtains platform information and uses an AI large model to reason about the platform information to obtain an execution instruction:
[0011] The Internet of Things platform conducts user confirmation based on the platform information, establishes an independent database, and transfers the multi-modal feature parameters to the independent database to obtain cloud data;
[0012] The Internet of Things platform establishes corresponding security verification information according to the execution instruction, sends it to the mobile device, and conducts security authentication based on the mobile device and the security verification information. The mobile device obtains the access address of the cloud data;
[0013] The mobile device accesses the cloud data according to the access address of the cloud data, obtains the device information of the corresponding intelligent device, and establishes a connection between the mobile device and the intelligent device to perform the required operations.
[0014] In an optional implementation manner, the intelligent device performs semantic recognition and natural language understanding on the voice information to obtain the user intention, including:
[0015] The intelligent device processes the voice information into text features, analyzes and constructs a graph structure through grammar unit analysis, and uses a page ranking algorithm to evaluate the importance of nodes according to the graph structure to obtain importance weights;
[0016] Using natural language processing technology, map the text features into vector form to obtain vector features reflecting semantic relationships;
[0017] According to the importance weights and the vector features, perform semantic classification through a label propagation algorithm, and obtain the user intention according to the classification result.
[0018] In an optional implementation manner, perform text analysis according to the user intention, the user identity information, and the usage scenario to obtain an operation command, and extract multi-modal feature parameters according to the operation command, including:
[0019] The intelligent device performs frame processing on the voice information, and obtains the voice feature of each frame by analyzing the amplitude, frequency, and waveform of each frame of voice information;
[0020] According to the voice feature, combined with a preset user identity model, authenticate the user to obtain user identity information;
[0021] According to the user intention, extract the usage scenario of the user through semantic analysis.
[0022] In an alternative embodiment, the intelligent device performs text analysis based on the user intention, the user identity information, and the usage scenario to obtain an operation command, including:
[0023] The intelligent device performs semantic analysis on the text using a natural language processing algorithm according to the user intention to obtain an initial command;
[0024] According to the initial command, using a graph matching algorithm, the closest operation command is matched from the existing knowledge graph to generate a candidate set of operation commands;
[0025] The operation commands in the candidate set of operation commands that do not conform to the user identity information and the usage scenario are eliminated to obtain a filtered candidate set of commands;
[0026] According to the filtered candidate set of commands, optimization processing is performed using a feature learning algorithm to obtain an operation command.
[0027] In an alternative embodiment, the intelligent device uploads the operation command and the multi-modal feature parameters to the Internet of Things platform. The Internet of Things platform obtains platform information and performs reasoning on the platform information using an AI large model to obtain an execution instruction, including:
[0028] The intelligent device obtains feature parameters based on a preset classification model according to the operation command and the multi-modal feature parameters, and then classifies the feature parameters using a classification neural network model to obtain a classification result;
[0029] Feature extraction and feature matching are performed according to the classification result, and the similarity is calculated. Whether to upload repeatedly is judged according to whether the similarity exceeds a preset similarity threshold;
[0030] If it is a repeated upload, the upload is rejected;
[0031] If there is no repeated upload, the operation command and the multi-modal feature parameters are saved to the Internet of Things platform as platform information;
[0032] The Internet of Things platform uses the platform information as input and performs reasoning using an AI large model, and the result is used as the execution instruction.
[0033] In an alternative embodiment, the Internet of Things platform performs user confirmation according to the platform information, establishes an independent database, and transfers the multi-modal feature parameters to the independent database to obtain cloud data, including:
[0034] The Internet of Things platform performs semantic understanding and analysis according to the platform information, confirms the user identity, and establishes a historical relationship graph of the interaction between the user and the system;
[0035] According to the historical relationship diagram, store relevant multi-modal feature parameters using a graph database structure, and create an independent database;
[0036] Use a query language to perform data query and update operations on the independent database, and upload the sorted multi-modal feature parameters to cloud storage to obtain cloud data.
[0037] In an optional implementation manner, the security authentication based on the mobile device and the security verification information to obtain the access address of the cloud data includes:
[0038] The Internet of Things platform generates a random code according to the security verification information, binds the random code with the security verification information to obtain a verification database;
[0039] Perform security verification based on the mobile device and the verification database to obtain an information verification result;
[0040] Make a judgment according to the information verification result. If the verification is passed, distribute the access address of the cloud data to the mobile device. If the verification fails, reject the distribution.
[0041] In an optional implementation manner, the security verification based on the mobile device and the verification database to obtain an information verification result includes:
[0042] Obtain the number of verification times according to the verification database;
[0043] If the number of verification times has not reached the preset number of verification times, pass and obtain security verification data;
[0044] If the number of verification times has reached the preset number of verification times, fail and abandon the verification;
[0045] Compare the security verification data with the random code in the verification database to obtain an information verification result.
[0046] In an optional implementation manner, the access to the cloud data according to the access address of the cloud data to obtain the device information of the corresponding intelligent device includes:
[0047] The mobile device obtains the call address for connecting to the independent database according to the access address of the cloud data;
[0048] By parsing the call address, determine the specific data table and fields in the independent database that store the intelligent device information as the parsing result;
[0049] Access the independent database according to the parsing result to obtain the device information of the corresponding intelligent device.
[0050] Second aspect, the present invention provides an AI voice processing method system based on intelligent devices, including:
[0051] A voice acquisition module, configured to acquire voice information uploaded by a user;
[0052] An intent extraction module, configured to perform semantic recognition and natural language understanding based on the voice information to obtain the user's intent;
[0053] An identity scenario extraction module, configured for the intelligent device to authenticate the user and extract the user's usage scenario based on the user's intent and the voice information, to obtain the user identity information and the usage scenario;
[0054] A data analysis module, configured to perform text analysis based on the user's intent, the user identity information, and the usage scenario to obtain an operation command, and extract multi-modal feature parameters according to the operation command;
[0055] A large model inference module, configured for the intelligent device to upload the operation command and the multi-modal feature parameters to the Internet of Things platform, the Internet of Things platform obtains platform information, and uses an AI large model to perform inference on the platform information to obtain an execution instruction:
[0056] A database establishment module, configured for the Internet of Things platform to confirm the user according to the platform information, establish an independent database, and transfer the multi-modal feature parameters to the independent database to obtain cloud data;
[0057] A security authentication module, configured to establish corresponding security verification information according to the execution instruction, send it to the mobile device, and perform security authentication based on the mobile device and the security verification information to obtain the access address of the cloud data;
[0058] An operation execution module, configured to access the cloud data according to the access address of the cloud data, obtain the device information of the corresponding intelligent device, and establish a connection between the mobile device and the intelligent device to execute the required operation.
[0059] Third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, where when the processor executes the computer program, it implements the AI voice processing method based on intelligent devices described in any one of the above.
[0060] Fourth aspect, the present invention further provides a computer-readable storage medium, where the computer-readable storage medium includes a stored computer program, and when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the AI voice processing method based on intelligent devices described in any one of the above.
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] The present invention discloses an AI voice processing method based on intelligent devices, including obtaining voice information uploaded by a user; performing semantic recognition and natural language understanding according to the voice information to obtain the user's intention; performing text analysis according to the user's intention, and making a refined distinction of the user's identity and usage scenario to obtain an operation command, and extracting multi-modal feature parameters according to the operation command; the intelligent device uploads the operation command and the multi-modal feature parameters to the Internet of Things platform, the Internet of Things platform obtains platform information, and uses an AI large model to perform reasoning on the platform information to obtain an execution instruction: the Internet of Things platform performs user confirmation according to the platform information, establishes an independent database, transfers the multi-modal feature parameters to the independent database to obtain cloud data; establishes corresponding security verification information according to the execution instruction, sends it to the mobile device, performs security authentication according to the mobile device and the security verification information to obtain the access address of the cloud data; accesses the cloud data according to the access address of the cloud data, obtains the device information of the corresponding intelligent device, and establishes a connection between the mobile device and the intelligent device to perform the required operation.
[0063] Through the processing of the voice information uploaded by the user, the present invention can realize more accurate and intelligent device control and data interaction. In the specific implementation process, the system first extracts the user's intention through voice recognition and natural language understanding technologies, generates specific operation commands according to the intention, and then analyzes the required data. This data is uploaded to the Internet of Things platform after being processed, and is inferred and judged by an AI large model to form the final execution instruction. These execution instructions will be transferred to an independent database for storage and management after user confirmation to ensure the security and independence of the data. Then, the system generates and sends corresponding security verification information through the mobile device to perform strict identity verification to ensure the access rights of the cloud data. After the user is successfully authenticated, the specific information related to the intelligent device can be obtained by accessing the specified cloud data, and an effective connection is established between the mobile device and the intelligent device to achieve precise control and operation. The present invention can efficiently and accurately identify the user's intention, improve the convenience and intelligence level of the interaction, and greatly improve the recognition accuracy of the voice processing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic flowchart of an AI voice processing method based on intelligent devices provided by the first embodiment of the present invention;
[0065] Figure 2It is a schematic structural diagram of an AI voice processing system based on an intelligent device provided by the second embodiment of the present invention. Detailed implementation manners
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] Referring to Figure 1 , the first embodiment of the present invention provides an AI voice processing method based on an intelligent device, including the following steps:
[0068] S11, the intelligent device obtains the voice information uploaded by the user;
[0069] S12, the intelligent device performs semantic recognition and natural language understanding according to the voice information to obtain the user's intention;
[0070] S13, the intelligent device authenticates the user according to the user's intention and the voice information and extracts the user's usage scenario to obtain the user identity information and the usage scenario;
[0071] S14, perform text analysis according to the user's intention, the user identity information and the usage scenario to obtain an operation command, and extract multi-modal feature parameters according to the operation command;
[0072] S15, the intelligent device uploads the operation command and the multi-modal feature parameters to the Internet of Things platform. The Internet of Things platform obtains platform information and uses an AI large model to reason about the platform information to obtain an execution instruction:
[0073] S16, the Internet of Things platform performs user confirmation according to the platform information, establishes an independent database, and transfers the multi-modal feature parameters to the independent database to obtain cloud data;
[0074] S17, the Internet of Things platform establishes corresponding security verification information according to the execution instruction, sends it to the mobile device, and performs security authentication according to the mobile device and the security verification information to obtain the access address of the cloud data;
[0075] S18, the mobile device accesses the cloud data according to the access address of the cloud data, obtains the device information of the corresponding intelligent device, and establishes a connection between the mobile device and the intelligent device to perform the required operations.
[0076] In step S11, the intelligent device needs to obtain the voice information uploaded by the user.
[0077] It should be noted that the user's voice information is input through intelligent devices (such as intelligent speakers, smartphones, smart home devices, etc.). These devices are built with voice recognition functions that can capture the user's sound signals in real time and convert them into digital voice data. The process of obtaining voice information includes the acquisition of sound signals, the transmission of sound data, and the preliminary processing of voice content (such as noise reduction, echo cancellation, etc.) to ensure that the system can receive user instructions efficiently and accurately in various environments (such as noisy places). The acquisition of this voice information is crucial for subsequent operations because it contains the user's true intentions and needs. Through accurate voice data acquisition, the system can accurately parse the user's instructions and lay a foundation for subsequent semantic recognition and natural language understanding.
[0078] In step S12, the intelligent device performs semantic recognition and natural language understanding based on the voice information to obtain the user's intention.
[0079] In one implementation, the performing semantic recognition and natural language understanding based on the voice information to obtain the user's intention includes:
[0080] The intelligent device processes the voice information into text features, analyzes and constructs a graph structure through syntactic units, and based on the graph structure, uses the PageRank algorithm to evaluate the importance of nodes to obtain importance weights;
[0081] Using natural language processing technology, maps the text features into vector form to obtain vector features reflecting semantic relationships;
[0082] Based on the importance weights and the vector features, performs semantic classification through the label propagation algorithm and obtains the user's intention according to the classification result.
[0083] It should be noted that the PageRank algorithm is used to evaluate the importance of each node in the text graph structure. Its core idea is to evaluate the importance of each node according to the relationships and link strengths between nodes. In the process of semantic recognition, the PageRank algorithm can determine which words or phrases are more crucial in the user's intention by calculating the relative importance between syntactic units. For example, in the sentence "Increase the air conditioner temperature", the algorithm will evaluate the critical weight of the verb "increase" by analyzing the relationship between "increase", "air conditioner", and "temperature", helping the system accurately understand the user's instruction.
[0084] It should be noted that vector features are the result of mapping text features into a vector space through natural language processing techniques, which are achieved through word embedding techniques. Vector features can reflect the semantic relationships between words, allowing the system to better understand the semantic similarities and differences between words. When processing user speech, vector features can convert each word or phrase into a high-dimensional vector, where words with similar meanings will have similar vector representations. This vectorized representation helps capture the context information and semantic relationships of the sentence, enabling the system to accurately identify the user's intention. In this way, the system can not only understand the meaning of a single word but also the semantics of the word in a specific context, thereby accurately obtaining the user's needs.
[0085] In step S13, the intelligent device authenticates the user according to the user intention and the speech information, and extracts the user usage scenario to obtain the user identity information and the usage scenario.
[0086] In one implementation, the intelligent device authenticates the user according to the user intention and the speech information, and extracts the user usage scenario to obtain the user identity information and the usage scenario, including:
[0087] The intelligent device performs frame splitting on the speech information, and obtains the speech features of each frame by analyzing the amplitude, frequency, and waveform of each frame of speech information;
[0088] According to the speech features, combined with a preset user identity model, the user is authenticated to obtain the user identity information;
[0089] According to the user intention, the user's usage scenario is extracted through semantic analysis.
[0090] It should be noted that frame splitting is because the speech signal is continuous. Frame splitting is to divide it into multiple short-time segments (frames), and each frame usually contains 20 - 30 ms of speech data, approximately the time of one word or phrase. Frame splitting is beneficial to extracting the features of each word spoken by the speaker, so as to more accurately obtain the user's identity. For example, assuming a speech signal lasts for 1 second, and frame splitting is performed with a frame length of 25 ms and a frame shift of 10 ms, about 75 frames of speech data can be obtained.
[0091] It should be noted that amplitude reflects the strength of the speech signal. For example, a high amplitude corresponds to the user speaking loudly, which can be used to judge the user's mood or environmental noise. Frequency represents the vibration speed of the speech signal and is related to the pitch. For example, the frequency of male speech is usually lower than that of female speech. High frequency may correspond to clear vowels, and low frequency may correspond to voiced sounds. The waveform describes the shape of the speech signal and reflects the periodicity and regularity of the speech. For example, a regular waveform corresponds to a stable vowel, and an irregular waveform corresponds to a consonant or noise.
[0092] It should be noted that identity verification is performed by combining voice features and a preset user identity model (such as a voiceprint recognition model based on deep learning). Assume that the voiceprint model of user A has been registered. When user A speaks, their voice features are extracted and compared with the model. If the feature matching degree is higher than a threshold (such as 90%), the verification passes, and the identity information of user A is obtained.
[0093] It should be noted that the usage scenario is extracted from the user's intention through semantic analysis. For example, when the user says "play light music in the kitchen", the text "play light music in the kitchen" is obtained after speech recognition. Semantic analysis identifies the keywords "kitchen" and "light music", thereby extracting the usage scenario as "the music playing requirement in the kitchen scenario".
[0094] In step S14, according to the user intention, the user identity information, and the usage scenario, text analysis is performed to obtain an operation command, and multi-modal feature parameters are extracted according to the operation command.
[0095] It should be noted that the multi-modal feature parameters refer to a structural parameter composed of an intelligent device identifier (such as an intelligent air conditioner ID), an operation type (such as temperature adjustment), a numerical parameter (such as a target temperature value of 25°C), operation permission information, and current environmental parameters. It can reflect the relevant characteristics of the operation instruction and facilitate the management of the operation instruction by the Internet of Things platform. Exemplarily, the operation command is a complete sentence. For example: Air conditioner 1 receives a command from a young person to adjust the indoor temperature to 20°C under the condition of a room temperature of 30°C; through semantic segmentation, the required keywords are extracted from the operation command, such as: young person, room temperature of 30 degrees Celsius, adjust the indoor temperature to 20°C, and the young person has full operation permission for air conditioner 1.
[0096] In one implementation, the intelligent device performs text analysis to obtain an operation command according to the user intention, the user identity information, and the usage scenario, including:
[0097] The intelligent device performs semantic analysis on the text using a natural language processing algorithm according to the user intention to obtain an initial command;
[0098] According to the initial command, the graph matching algorithm is used to match the closest operation command from the existing knowledge graph to generate a candidate set of operation commands;
[0099] The operation commands that do not conform to the user identity information and the usage scenario in the candidate set of operation commands are eliminated to obtain a filtered candidate set of operation commands;
[0100] According to the filtered candidate set of operation commands, optimization processing is performed using a feature learning algorithm to obtain an operation command.
[0101] It should be noted that the graph matching algorithm is a technology used to find the operation command closest to the user's input in the knowledge graph. By calculating the similarity between nodes and edges in the graph, the graph matching algorithm can identify the correspondence between the user's intention and the existing operation commands in the knowledge graph. In practical applications, the knowledge graph contains a large number of entities and relationships. Based on this structured data, the graph matching algorithm can quickly find the operation command that best matches the user's input intention. For example, when the user inputs "increase the air conditioner temperature", the graph matching algorithm will match relevant action commands in the graph according to the semantic features of "increase" and "air conditioner temperature", such as "increase temperature" or "raise the air conditioner temperature".
[0102] It is worth noting that the feature learning algorithm is a data-driven method used to optimize the operation command candidate set. Feature learning extracts features from a large amount of historical data and trains a model based on these features to evaluate the matching degree of each candidate command with the user's needs. In the present invention, the feature learning algorithm not only considers the semantic consistency of the candidate commands, but also combines context information and the user's historical behavior to further optimize the selection of operation commands. Feature learning can help the system more accurately understand the user's needs under complex language inputs, so as to select the operation command that best meets the user's intention. Exemplarily, if the user repeatedly selects to adjust the air conditioner temperature, the system will identify this preference through the feature learning model and preferentially select relevant operation commands, thereby providing more personalized and accurate services.
[0103] It is worth noting that in order to obtain the filtered command candidate set, it is first necessary to filter the operation command candidate set according to the user identity information and usage scenario. For example, assume that the user is an elderly person, and their usage scenario is to play music using smart devices at home. The operation command candidate set may contain various commands, such as "play music", "query stock information", "set alarm", etc. Since the elderly are generally not interested in stock information, and in the current music-playing scenario, the "query stock information" command does not match the user identity and usage scenario, so it is excluded. At the same time, although "set alarm" is related to the user identity, it also does not meet the user's immediate needs in the current music-playing scenario, so it will also be excluded. Finally, by excluding the operation commands that do not match the user identity information and usage scenario, the filtered command candidate set is obtained, such as "play music", "pause play", "adjust volume", etc., which are more in line with the actual needs of the elderly in the music-playing scenario.
[0104] In step S15, the smart device uploads the operation command and the multi-modal feature parameters to the Internet of Things platform. The Internet of Things platform obtains the platform information and uses the AI large model to reason about the platform information to obtain the execution instruction.
[0105] In one implementation, the smart device uploads the operation command and the multimodal feature parameters to the Internet of Things platform. The Internet of Things platform obtains platform information and uses an AI large model to reason about the platform information to obtain an execution instruction, including:
[0106] Based on a preset classification model, the smart device obtains feature parameters according to the operation command and the multimodal feature parameters, and then uses a classification neural network model to classify the feature parameters to obtain a classification result;
[0107] Feature extraction and feature matching are performed according to the classification result, and the similarity is calculated. Whether to repeat the upload is judged according to whether the similarity exceeds a preset similarity threshold;
[0108] If it is a repeated upload, the upload is rejected;
[0109] If there is no repeated upload, the operation command and the multimodal feature parameters are saved to the Internet of Things platform as platform information;
[0110] The Internet of Things platform uses the platform information as input and uses an AI large model for reasoning, and the obtained result is used as the execution instruction.
[0111] It should be noted that the platform information refers to the combination of the operation command and the multimodal feature parameters.
[0112] It should be noted that the preset classification model can extract key feature parameters from the uploaded operation command and multimodal feature parameters based on historical data and previously defined features. These feature parameters represent the specific details of the user's intention, such as the target device, operation type, data range, etc. Through the preset classification model, the system can ensure appropriate feature extraction for different types of commands and data, so as to provide effective input for subsequent reasoning steps.
[0113] It is worth noting that the classification neural network model is used to further classify the extracted feature parameters. The neural network model can handle complex non-linear relationships and improve the classification accuracy through training and optimization. In this process, the model will identify which features are the most discriminative in a given command and classify them into different operation categories. Next, the system will perform feature extraction and feature matching according to the classification result, and calculate the similarity between different operation commands and existing operations. Through this similarity calculation, the system can judge whether the current operation is repeated with the previously uploaded command. If the similarity exceeds the preset threshold, the system will reject the repeated upload to avoid data redundancy or conflict; otherwise, if there is no repetition, the system will execute the upload and save the operation command and multimodal feature parameters to the Internet of Things platform as platform information for subsequent processing.
[0114] It should be noted that the AI large model refers to an artificial intelligence model trained with large-scale data and optimized by deep learning. It has powerful reasoning capabilities and can identify potential rules and patterns from complex data. In step S15, the AI large model generates execution instructions by analyzing the information on the Internet of Things platform and performing intelligent reasoning on various environmental factors and user inputs. Different from traditional models, the AI large model contains billions or even more parameters, can process more complex inputs, and has strong generalization and self-learning capabilities.
[0115] It is worth noting that the reasoning process of the AI large model can combine multiple data sources and context information to help the system make more accurate decisions in different situations. Exemplarily, the AI large model not only analyzes the operation commands and multi-modal feature parameters uploaded by users, but also infers the most suitable execution instructions for the current situation based on the current state of the device, historical operation records, and data of other Internet of Things devices. Through such a reasoning process, the AI large model can significantly improve the intelligence level of the system, making the execution instructions more in line with user needs and adaptable to complex and changing environmental conditions. In addition, the introduction of the AI large model can also support continuous self-optimization and update, continuously acquire knowledge from new data, and improve the overall performance and adaptability of the system.
[0116] In step S16, the Internet of Things platform performs user confirmation based on the platform information, establishes an independent database, transfers the multi-modal feature parameters to the independent database, and obtains cloud data.
[0117] In one implementation, the Internet of Things platform performs user confirmation based on the platform information, establishes an independent database, transfers the multi-modal feature parameters to the independent database, and obtains cloud data, including:
[0118] The Internet of Things platform performs semantic understanding and analysis based on the platform information, confirms the user identity, and establishes a historical relationship graph of the interaction between the user and the system;
[0119] According to the historical relationship graph, store the relevant multi-modal feature parameters in the structure of a graph database to create an independent database;
[0120] Use a query language to perform data query and update operations on the independent database, upload the sorted multi-modal feature parameters to cloud storage, and obtain cloud data.
[0121] It should be noted that the process of confirming the user's identity involves in-depth mining of the user's identity information, including comprehensive comparison through device, account information, and historical interaction data in the platform to ensure that the current operation is initiated by an authorized user. This identity confirmation not only relies on traditional usernames and passwords but also enhances security through device fingerprints to prevent unauthorized access.
[0122] It should be noted that establishing a historical relationship graph of the interaction between the user and the system provides a basis for subsequent data management and intelligent decision-making. This historical relationship graph not only records the user's operation history but also reflects the interaction patterns, preferences, and habits between the user and the system. The system uses a graph database structure to store this relationship information, and the graph database can efficiently store and query complex relationship data, supporting fast graph traversal and pattern matching. Through this structure, the system can accurately infer the user's needs in subsequent operations, improving response efficiency and operation accuracy.
[0123] It is worth noting that cloud storage can not only ensure the persistence and reliability of data but also provide efficient data support for subsequent intelligent reasoning and analysis. Through this process, the user's multi-modal feature parameters are efficiently and securely managed and can be seamlessly accessed between different devices and systems, enhancing the intelligence and user experience of the entire Internet of Things system.
[0124] In step S17, corresponding security verification information is established according to the execution instruction and sent to the mobile device. Security authentication is performed based on the mobile device and the security verification information to obtain the access address of the cloud data.
[0125] In one implementation, the security authentication based on the mobile device and the security verification information to obtain the access address of the cloud data includes:
[0126] The Internet of Things platform generates a random code according to the security verification information, binds the random code with the security verification information to obtain a verification database;
[0127] Security verification is performed based on the mobile device and the verification database to obtain an information verification result;
[0128] Judgment is made according to the information verification result. If the verification is passed, the access address of the cloud data is distributed to the mobile device. If the verification fails, the distribution is refused.
[0129] In one implementation, the security verification based on the mobile device and the verification database to obtain an information verification result includes:
[0130] The verification times are obtained according to the verification database;
[0131] If the number of verification attempts has not reached the preset number of verification attempts, it passes and obtains the security verification data;
[0132] If the number of verification attempts reaches the preset number of verification attempts, it fails and abandons the verification;
[0133] Compare the security verification data with the random code in the verification database to obtain the information verification result.
[0134] It should be noted that the information verification result is obtained by comparing the security verification data with the random code in the verification database. This process is mainly based on the principle of information security verification. By ensuring the uniqueness and timeliness of the verification information, it confirms the user's identity and prevents malicious attacks. The information verification principle is essentially a hash algorithm and encryption technology. During the verification process, the random code and the security verification information will undergo a hash process to generate an irreversible verification value. This verification value will be stored in the verification database. When the user attempts to verify, the system recalculates the verification value through the same hash algorithm and compares it with the value in the database to ensure the integrity and security of the verification process. In addition, the system may also use encryption technology to ensure that the data during the transmission process is not tampered with, thus ensuring the reliability of the information verification.
[0135] It should be noted that generating a random code and binding it to the security verification information is to increase the security of the verification process. The generated random code changes in each authentication process, which can prevent the intercepted and misused security information reused by attackers. Storing the random code and the security verification information together in the verification database ensures that each security authentication has a unique and unpredictable verification information, thus effectively preventing potential security risks.
[0136] It is worth noting that when performing security verification based on the mobile device and the verification database, the system first obtains the number of verification attempts to ensure the rationality of the verification process and prevent brute-force cracking. By setting the preset number of verification attempts, the system can allow users to perform security verification within a certain number of times. If the verification fails to reach the set number of times, it will automatically abandon further verification to prevent malicious attackers from obtaining access rights by continuously trying. If the number of verification attempts has not reached the preset number, the system will continue the security verification and generate the final information verification result by comparing the random code in the verification database with the security verification data.
[0137] In step S18, according to the access address of the cloud data, access the cloud data, obtain the device information of the corresponding intelligent device, and establish a connection between the mobile device and the intelligent device to perform the required operations.
[0138] In one implementation, accessing the cloud data according to the access address of the cloud data of the mobile device to obtain the device information of the corresponding intelligent device includes:
[0139] The mobile device obtains the call address for connecting to the independent database according to the access address of the cloud data;
[0140] By parsing the call address, determine the specific data table and fields storing the intelligent device information in the independent database and use it as the parsing result;
[0141] Access the independent database according to the parsing result to obtain the device information of the corresponding intelligent device.
[0142] It should be noted that the call address for connecting to the independent database is the interface between the cloud storage system and the independent database. Through it, the system can interact with the database storing the intelligent device information. The call address is an identifier pointing to a specific location in the database, which can ensure that the system accurately accesses the required database.
[0143] It is worth noting that during the process of parsing the call address, the system will identify the specific data table and fields storing the intelligent device information in the independent database. The detailed information of each intelligent device (such as device model, status, configuration parameters, etc.) is stored in specific tables and fields of the database. The purpose of parsing the call address is to accurately locate this information. The parsing result provides a comprehensive understanding of the database structure, enabling the system to accurately retrieve and access the relevant data. After obtaining the parsing result, the system will initiate a database access request based on this information and successfully extract the detailed device information related to the intelligent device from the independent database.
[0144] The present invention discloses an AI voice processing method based on intelligent devices, which includes obtaining the voice information uploaded by users; performing semantic recognition and natural language understanding according to the voice information to obtain the user's intention; performing text analysis based on the user's intention to obtain an operation command, and extracting multi-modal feature parameters according to the operation command; the intelligent device uploads the operation command and the multi-modal feature parameters to the Internet of Things platform, the Internet of Things platform obtains platform information, and uses an AI large model to reason about the platform information to obtain an execution instruction: the Internet of Things platform performs user confirmation according to the platform information, establishes an independent database, transfers the multi-modal feature parameters to the independent database to obtain cloud data; establishes corresponding security verification information according to the execution instruction, sends it to the mobile device, performs security authentication according to the mobile device and the security verification information, and obtains the access address of the cloud data; accesses the cloud data according to the access address of the cloud data, obtains the device information of the corresponding intelligent device, and establishes a connection between the mobile device and the intelligent device to perform the required operations.
[0145] Through the processing of the voice information uploaded by users, the present invention can achieve more accurate and intelligent device control and data interaction. In the specific implementation process, the system first extracts the user's intention through voice recognition and natural language understanding technologies, and generates specific operation commands according to this intention, and then analyzes the required data. This data is uploaded to the Internet of Things platform after being processed, and is inferred and judged by an AI large model to form the final execution instruction. These execution instructions will be transferred to an independent database for storage and management after user confirmation to ensure the security and independence of the data. Then, the system generates and sends corresponding security verification information through the mobile device to perform strict identity verification to ensure the access rights of the cloud data. After the user is successfully authenticated, the specific information related to the intelligent device can be obtained by accessing the specified cloud data, and an effective connection is established between the mobile device and the intelligent device to achieve accurate control and operation. The present invention can efficiently and accurately identify the user's intention, improve the convenience and intelligence level of the interaction, and greatly improve the recognition accuracy of the voice processing system.
[0146] Referring to Figure 2 , the second embodiment of the present invention provides an AI voice processing system based on intelligent devices, including:
[0147] A voice acquisition module for obtaining the voice information uploaded by users;
[0148] An intention extraction module for performing semantic recognition and natural language understanding according to the voice information to obtain the user's intention;
[0149] An identity scenario extraction module, configured to authenticate the user and extract the user's usage scenario according to the user intention and the voice information by the intelligent device, so as to obtain user identity information and usage scenario;
[0150] A data analysis module, configured to perform text analysis according to the user intention, the user identity information and the usage scenario to obtain an operation command, and extract multi-modal feature parameters according to the operation command;
[0151] A large model inference module, configured to upload the operation command and the multi-modal feature parameters by the intelligent device to the Internet of Things platform, the Internet of Things platform obtains platform information, and uses an AI large model to infer the platform information to obtain an execution instruction:
[0152] A database establishment module, configured to confirm the user by the Internet of Things platform according to the platform information, establish an independent database, and transfer the multi-modal feature parameters to the independent database to obtain cloud data;
[0153] A security authentication module, configured to establish corresponding security verification information according to the execution instruction, send it to the mobile device, and perform security authentication according to the mobile device and the security verification information to obtain the access address of the cloud data;
[0154] An operation execution module, configured to access the cloud data according to the access address of the cloud data, obtain the device information of the corresponding intelligent device, and establish a connection between the mobile device and the intelligent device to perform the required operations.
[0155] It should be noted that an AI voice processing system based on an intelligent device provided in an embodiment of the present invention is used to execute all the process steps of an AI voice processing method based on an intelligent device in the above embodiment. The working principles and beneficial effects of the two correspond one by one, so they will not be elaborated here.
[0156] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a voice acquisition program. When the processor executes the computer program, it implements the steps in each of the above embodiments of the AI voice processing method based on an intelligent device, such as Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments, such as the voice acquisition module.
[0157] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0158] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than the above, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0159] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire electronic device through various interfaces and circuits.
[0160] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and by invoking the data stored in the memory, the processor implements various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0161] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or system, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0162] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the system embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0163] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An AI voice processing method based on intelligent devices, characterized in that: include: The smart device obtains the voice information uploaded by the user; The smart device performs semantic recognition and natural language understanding based on the voice information to obtain the user's intention; The smart device authenticates the user and extracts the user usage scenario according to the user intention and the voice information to obtain the user identity information and the usage scenario; The smart device performs text analysis according to the user intention to obtain an operation command candidate set, removes the operation commands that do not conform to the user identity information and the usage scenario from the operation command candidate set, obtains a target operation command, and extracts multimodal feature parameters according to the target operation command; The smart device uploads the operation command and the multimodal feature parameters to the IoT platform. The IoT platform obtains the platform information and uses the AI big model to infer the platform information to obtain the execution instruction: The Internet of Things platform performs user confirmation according to the platform information, establishes an independent database, transfers the multimodal feature parameters to the independent database, and obtains cloud data; The IoT platform establishes corresponding security verification information according to the execution instruction, sends it to the mobile device, performs security authentication according to the mobile device and the security verification information, and the mobile device obtains the access address of the cloud data; The mobile device accesses the cloud data according to the access address of the cloud data, obtains the device information of the corresponding smart device, establishes a connection between the mobile device and the smart device, and performs the required operation; The smart device performs semantic recognition and natural language understanding based on the voice information to obtain the user's intention, including: The intelligent device processes the voice information into text features, analyzes and constructs a graph structure through grammatical units, and evaluates the importance of nodes using a page sorting algorithm according to the graph structure to obtain importance weights; Using natural language processing technology, the text features are mapped into vector form to obtain vector features reflecting semantic relationships; According to the importance weight and the vector feature, semantic classification is performed through a label propagation algorithm, and the user intention is obtained according to the classification result.
2. The AI voice processing method based on intelligent device according to claim 1, characterized in that: The smart device authenticates the user and extracts the user usage scenario according to the user intention and the voice information to obtain the user identity information and the usage scenario, including: The intelligent device performs frame processing on the voice information, and obtains the voice features of each frame by analyzing the amplitude, frequency, and waveform of each frame of voice information; According to the voice features, combined with a preset user identity model, the user's identity is authenticated to obtain user identity information; According to the user intention, the user's usage scenario is extracted through semantic analysis.
3. The AI voice processing method based on intelligent device according to claim 1, characterized in that: The smart device performs text analysis according to the user intention to obtain an operation command candidate set, removes the operation commands in the operation command candidate set that do not conform to the user identity information and the usage scenario, and obtains the target operation command, including: The smart device performs semantic analysis on the text using a natural language processing algorithm according to the user's intention to obtain an initial command; According to the initial command, the graph matching algorithm is used to match the closest operation command from the existing knowledge graph to generate a candidate set of operation commands; Eliminate the operation commands that do not conform to the user identity information and the usage scenario from the operation command candidate set to obtain a screening command candidate set; According to the screening command candidate set, an optimization process is performed using a feature learning algorithm to obtain an operation command.
4. The AI voice processing method based on intelligent device according to claim 1, characterized in that: The smart device uploads the operation command and the multimodal feature parameters to the IoT platform, and the IoT platform obtains platform information, and uses the AI big model to infer the platform information to obtain execution instructions, including: The smart device obtains characteristic parameters according to the operation command and the multimodal characteristic parameters based on a preset classification model, and then classifies the characteristic parameters using a classification neural network model to obtain a classification result; Extracting and matching features according to the classification results, and calculating similarity, and judging whether to upload repeatedly according to whether the similarity exceeds a preset similarity threshold; If the upload is repeated, the upload will be rejected; If there is no repeated upload, the operation command and the multimodal feature parameters are saved to the Internet of Things platform as platform information; The IoT platform uses the platform information as input, uses the AI big model for reasoning, and obtains the result as the execution instruction.
5. The AI voice processing method based on intelligent device according to claim 1, characterized in that: The Internet of Things platform performs user confirmation according to the platform information, establishes an independent database, transfers the multimodal feature parameters to the independent database, and obtains cloud data, including: The IoT platform performs semantic understanding and analysis based on the platform information, confirms the user's identity, and establishes a historical relationship diagram of the user's interaction with the system; According to the historical relationship graph, a graph database structure is used to store relevant multimodal feature parameters to create an independent database; The independent database is queried and updated using a query language, and the sorted multimodal feature parameters are uploaded to a cloud storage to obtain cloud data.
6. The AI voice processing method based on intelligent device according to claim 1, characterized in that: The performing security authentication according to the mobile device and the security verification information to obtain the access address of the cloud data includes: The Internet of Things platform generates a random code according to the security verification information, binds the random code to the security verification information, and obtains a verification database; Perform security verification based on the mobile device and the verification database to obtain an information verification result; The judgment is made based on the information verification result. If the verification passes, the access address of the cloud data is distributed to the mobile device. If the verification fails, the distribution is refused.
7. The AI voice processing method based on intelligent device according to claim 6, characterized in that: The performing security verification according to the mobile device and the verification database to obtain an information verification result includes: According to the verification database, obtaining the number of verifications; If the verification times do not reach the preset verification times, the verification passes and the security verification data is obtained; If the verification times reach the preset number, the verification will be rejected and abandoned; The security verification data is compared with the random code of the verification database to obtain an information verification result.
8. The AI voice processing method based on intelligent device according to claim 1, characterized in that: The mobile device accesses the cloud data according to the access address of the cloud data, and obtains device information of the corresponding smart device, including: The mobile device obtains a calling address for connecting to an independent database according to the access address of the cloud data; By parsing the call address, determining the specific data table and field storing the smart device information in the independent database as the parsing result; The independent database is accessed according to the analysis result to obtain device information of the corresponding smart device.
9. An AI voice processing system based on smart devices, characterized in that: include: Voice acquisition module, used to obtain voice information uploaded by users; An intention extraction module is used to perform semantic recognition and natural language understanding based on the voice information to obtain the user's intention; An identity scenario extraction module is used for the smart device to authenticate the user and extract the user usage scenario according to the user intention and the voice information, so as to obtain the user identity information and the usage scenario; A data analysis module, configured to perform text analysis to obtain operation commands according to the user intention, the user identity information and the usage scenario, and extract multimodal feature parameters according to the operation commands; The large model inference module is used for the smart device to upload the operation command and the multimodal feature parameters to the IoT platform. The IoT platform obtains the platform information and uses the AI large model to infer the platform information to obtain the execution instruction: A database establishment module, used for the Internet of Things platform to perform user confirmation according to the platform information, establish an independent database, transfer the multimodal feature parameters to the independent database, and obtain cloud data; A security authentication module, used to establish corresponding security authentication information according to the execution instruction, send it to the mobile device, perform security authentication according to the mobile device and the security authentication information, and obtain the access address of the cloud data; An operation execution module, used to access the cloud data according to the access address of the cloud data, obtain the device information of the corresponding smart device, establish a connection between the mobile device and the smart device, and execute the required operation; The smart device performs semantic recognition and natural language understanding based on the voice information to obtain the user's intention, including: The intelligent device processes the voice information into text features, analyzes and constructs a graph structure through grammatical units, and evaluates the importance of nodes using a page sorting algorithm according to the graph structure to obtain importance weights; Using natural language processing technology, the text features are mapped into vector form to obtain vector features reflecting semantic relationships; According to the importance weight and the vector feature, semantic classification is performed through a label propagation algorithm, and the user intention is obtained according to the classification result.
Citation Information
Patent Citations
Intelligent monitoring method and system based on Internet of Things, medium and program product
CN119135741A
Internet of Things equipment linkage control method and system based on cloud computing, and storage medium
CN119472327A