Artificial intelligence voice interaction system based on big data

By using multi-dimensional data collection and multi-level technical means, we have solved the technical problems that were not solved in the existing technology, achieved efficient and personalized interaction, and improved the system's response speed and user experience.

CN121122268APending Publication Date: 2025-12-12BEIJING CHUANGLIAN YUNRUI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511275677.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-12-12

Smart Images

  • Figure CN121122268A_ABST
    Figure CN121122268A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and big data fusion, and particularly relates to an artificial intelligence voice interaction system based on big data, and the system comprises a multi-dimensional data collection module which is used for collecting user voice data, real-time scene data and equipment state data, and providing basic data support for subsequent processing; the big data preprocessing module is used for carrying out cleaning and feature extraction on the collected original data, removing invalid data and improving the subsequent analysis efficiency; and the multi-level semantic analysis module is used for associating a multi-instruction sequence through a sequential logic algorithm and adjusting an analysis priority in combination with scene data. Through progressive logic of the multi-dimensional data acquisition module, the big data preprocessing module and the multi-level semantic analysis module, the situation that original data containing noise and invalid fields are directly used for semantic analysis is avoided, and repeated instruction operation caused by analysis errors of a user can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and big data fusion technology, specifically to an artificial intelligence voice interaction system based on big data. Background Technology

[0002] With the rapid development of artificial intelligence and big data technologies, voice interaction has become one of the core methods of human-computer interaction. Currently, AI-powered voice interaction systems based on big data collect user voice data, scene data, and device status data, and combine them with AI algorithms to achieve command recognition and execution. These systems are widely used in fields such as homes, transportation, and healthcare. For example, in smart homes, users can control devices such as lights, air conditioners, and audio systems using voice commands; in in-vehicle terminals, users can use voice commands to check traffic conditions, play music, and make phone calls.

[0003] However, users' demands for voice interaction are evolving from "single command response" to "multi-command collaboration in complex scenarios," "personalized adaptation," and "low-latency feedback." Existing systems rely on big data to support basic interaction, but in practical applications, deficiencies in semantic understanding logic, data processing architecture, and user profile adaptation make it difficult to meet the demands for high-quality interaction, resulting in several unresolved technical shortcomings; these are detailed below:

[0004] (i) Insufficient multi-instruction correlation and scenario adaptability in semantic understanding

[0005] Existing systems often employ a parsing logic of "keyword extraction + single command matching," which fails to effectively connect multiple operational requirements within complex commands and does not dynamically adjust parsing priorities based on real-time scenarios. For example, in a smart home scenario, if a user gives the command "first set the bedroom air conditioner to 26 degrees, then turn off the study light, and finally play soothing music in the bedroom," existing systems, unable to recognize the "first / then / last" command sequence logic, only execute the single command "set the bedroom air conditioner to 26 degrees," or confuse the operation objects (mistakenly turning off the bedroom light or turning on the study speaker), causing the user to have to issue the command 3-4 times to complete the task, reducing interaction efficiency by more than 60%.

[0006] (ii) Mismatch between big data processing and real-time requirements

[0007] Existing systems largely rely on a "full-data cloud processing" model, where all voice and scene data are uploaded to the cloud for analysis and computation. When faced with scenarios involving concurrent processing of multi-dimensional data, this model is prone to response delays due to data transmission latency and excessive cloud computing power consumption. For example, in a vehicle scenario, if a user issues a command such as "find a gas station within 5 kilometers ahead, avoid congested areas, and announce the remaining fuel level for the next mileage," the existing system must simultaneously process large amounts of map data (road conditions, gas station locations) and vehicle status data (fuel level, speed). Because full-data processing relies on cloud computing, the response delay exceeds 2 seconds, and can reach over 5 seconds at high speeds or when network signals are weak. This not only affects user experience but also poses a driving safety hazard.

[0008] (III) Lack of personalized interaction adaptation capabilities

[0009] The existing system lacks a dynamically updated user profile model, making it unable to adapt interaction parameters based on different user characteristics (age, gender, usage habits). This results in inconsistent user experiences across different users within the same system. For example, within the same household, an elderly person might give the command, "Help, turn up the TV volume and maximize the font," while a teenager might give the command, "Help, turn down the TV volume and switch to game mode." Because the existing system cannot identify users through voice characteristics and historical interaction data, users must repeatedly give personalized commands each time. This is especially problematic for elderly users unfamiliar with the system, requiring them to repeatedly adjust voice and font parameters, creating a high barrier to entry. Statistics show that elderly users take 3-5 more steps and spend 2-3 times more time using the existing system to fulfill personalized needs compared to teenagers.

[0010] Based on the above, an artificial intelligence voice interaction system based on big data is invented. Summary of the Invention

[0011] To address the aforementioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution:

[0012] An AI-powered voice interaction system based on big data, comprising:

[0013] The multi-dimensional data acquisition module is used to collect user voice data, real-time scene data, and device status data, providing basic data support for subsequent processing;

[0014] The big data preprocessing module is used to clean and extract features from the collected raw data, remove invalid data, and improve the efficiency of subsequent parsing.

[0015] The multi-level semantic parsing module is used to associate the order of multiple instructions through temporal logic algorithms and adjust the parsing priority based on scenario data;

[0016] The user profile dynamic update module is used to update user preferences in real time through collaborative filtering algorithms to adapt to personalized interaction needs.

[0017] The multi-device collaboration conflict handling module is used to resolve operational conflicts and mutual exclusion conflicts of device parameters when multiple devices respond to the same command at the same time.

[0018] The edge-cloud collaborative processing module is used to dynamically allocate processing nodes according to network status, with simple instructions processed at the edge and complex big data tasks processed in the cloud;

[0019] The intelligent decision-making and feedback module is used to generate device control commands based on semantic parsing results and user profiles, and to provide feedback on the execution results to the user.

[0020] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the multi-dimensional data acquisition module includes:

[0021] The voice acquisition unit is used to acquire the voiceprint features and intonation features of the user's voice commands using a noise-canceling microphone array;

[0022] The scene data acquisition unit is used to acquire scene parameters through temperature sensors, position sensors, and a time module.

[0023] The device status acquisition unit is used to establish a communication connection with associated devices and obtain the current operating status of the devices in real time.

[0024] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the big data preprocessing module includes:

[0025] The data cleaning unit is used to correct recognition errors in voice commands through a speech recognition error correction algorithm.

[0026] The feature extraction unit is used to extract instruction keywords and user voiceprint features from speech data, and time-scene correlation features from scene data.

[0027] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the multi-level semantic parsing module includes:

[0028] Basic semantic units are used to parse the operands and parameters of a single instruction.

[0029] Context semantic unit, used to identify instruction order and associate dependencies between multiple instructions through timing logic algorithms;

[0030] Scene semantic unit, used to adjust parsing priority based on scene data.

[0031] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the user profile dynamic update module includes:

[0032] A user profile building unit is used to create a user profile based on static and dynamic data.

[0033] The update logic unit is used to synchronize the current command parameters and user feedback to the profile after each interaction, and optimize preference recommendations through collaborative filtering algorithms.

[0034] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the multi-device collaborative conflict handling module includes:

[0035] The conflict detection unit is used to first identify the set of devices and operating parameters involved in the instruction, and then determine the conflict through the device association matrix; then it performs logical verification on the device parameter conflicts.

[0036] The conflict resolution unit prioritizes the device closest to the user's current location when dealing with multi-device response conflicts. If the location cannot be determined, it selects the device with the highest historical usage frequency based on the user profile. For parameter mutual exclusion conflicts, it makes decisions based on scenario priority.

[0037] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the edge-cloud collaborative processing module includes:

[0038] Edge units are used to process simple device control commands;

[0039] The cloud unit is used to process complex tasks that require big data support;

[0040] The dynamic allocation unit is used to adjust the processing nodes based on the network status. When the network is weak, high-frequency instructions are cached at the edge to reduce reliance on the cloud.

[0041] As a preferred embodiment of the big data-based artificial intelligence voice interaction system described in this invention, the intelligent decision-making and feedback module includes:

[0042] The decision generation unit is used to send precise control commands to associated devices;

[0043] The voice feedback unit is used to broadcast the execution results to the user through TTS speech synthesis technology;

[0044] The exception handling unit is used to provide real-time feedback of exception information when the device is unresponsive.

[0045] Compared with existing technologies:

[0046] I. Consistent data flow significantly reduces semantic parsing errors.

[0047] Through the progressive logic of multi-dimensional data acquisition module, big data preprocessing module and multi-level semantic parsing module, it avoids directly using raw data containing noisy or invalid fields for semantic parsing, thereby reducing the number of repeated command operations caused by parsing errors by users.

[0048] II. Prioritizing Dependencies to Ensure Personalized Interactions Are Implemented

[0049] By setting up a multi-level semantic parsing module, a user profile dynamic update module, and a multi-device collaborative conflict handling module, it ensures that the latest user preference data can be called when resolving conflicts, avoiding unfounded random decisions, and is suitable for groups such as the elderly and children who are sensitive to the complexity of operation.

[0050] III. Balancing real-time performance and accuracy to shorten interactive response latency

[0051] By setting up a multi-device collaborative conflict handling module, an edge-cloud collaborative processing module, and an intelligent decision-making and feedback module, it can resolve device conflicts before allocating processing nodes, avoiding "secondary rework" caused by conflicting instructions at the edge / cloud; at the same time, it allocates nodes according to the complexity of instructions, taking into account both real-time performance and data support requirements.

[0052] IV. Precise resource allocation reduces unnecessary computing power consumption in the system.

[0053] By "preprocessing and filtering invalid data" and "removing contradictory instructions through conflict handling", the system ensures that subsequent edge-cloud collaborative processing targets only "high-quality, conflict-free" valid instructions, avoiding wasted computing power, making the system run more lightweight, and adapting to resource-constrained scenarios such as smart wearable devices.

[0054] V. Complete interactive loop enhances the continuity of user experience.

[0055] By forming a seamless interactive chain from data acquisition to final voice feedback, the output of each module becomes the precise input of the next module, avoiding interaction lag caused by "data interruption" and "decision-making gap". Users do not need to wait or repeat operations, thus improving the continuity of interaction. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of the overall framework of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0058] This invention provides an artificial intelligence voice interaction system based on big data. Please refer to [link / reference]. Figure 1 ,include:

[0059] The multi-dimensional data acquisition module is used to collect user voice data, real-time scene data, and device status data, providing basic data support for subsequent processing;

[0060] The big data preprocessing module is used to clean and extract features from the collected raw data, remove invalid data, and improve the efficiency of subsequent parsing.

[0061] The multi-level semantic parsing module is used to associate the order of multiple instructions through temporal logic algorithms and adjust the parsing priority based on scenario data;

[0062] The user profile dynamic update module is used to update user preferences in real time through collaborative filtering algorithms to adapt to personalized interaction needs.

[0063] The multi-device collaboration conflict handling module is used to resolve operational conflicts and mutual exclusion conflicts of device parameters when multiple devices respond to the same command at the same time.

[0064] The edge-cloud collaborative processing module is used to dynamically allocate processing nodes according to network status, with simple instructions processed at the edge and complex big data tasks processed in the cloud;

[0065] The intelligent decision-making and feedback module is used to generate device control commands based on semantic parsing results and user profiles, and to provide feedback on the execution results to the user.

[0066] The multi-dimensional data acquisition module includes:

[0067] The voice acquisition unit uses a noise-canceling microphone array to filter out environmental noise (such as water flow noise and traffic noise) and acquire the voiceprint and intonation features of the user's voice commands.

[0068] The scene data acquisition unit is used to acquire scene parameters (such as current time, user location, and ambient temperature) through temperature sensors, position sensors, and time modules.

[0069] The equipment status acquisition unit is used to establish communication connections with associated equipment (air conditioning, lights, television, vehicle ECU, etc.) to obtain the current operating status of the equipment in real time (such as air conditioning temperature, light on / off status, vehicle fuel level).

[0070] The big data preprocessing module includes:

[0071] The data cleaning unit is used to correct recognition errors in voice commands using a speech recognition error correction algorithm (such as correcting the error based on context when "gentle piano music" is misrecognized as "light and heavy piano music").

[0072] The feature extraction unit is used to extract instruction keywords (such as "air conditioner 26 degrees" and "turn off the lights"), user voiceprint features, and time-scene association features (such as "20:00-bedroom-sleeping scene") from voice data.

[0073] The multi-level semantic parsing module includes:

[0074] Basic semantic units are used to parse the operation objects and parameters of a single instruction (e.g., "bedroom air conditioner" as the object and "26 degrees" as the parameter);

[0075] Context semantic units are used to identify the order of instructions (such as "first / then / last") through temporal logic algorithms and to associate the dependencies between multiple instructions (such as "play sleep music" must be executed after "adjust the air conditioner and turn off the lights").

[0076] Scene semantic units are used to adjust the parsing priority based on scene data (e.g., in a vehicle scene, "query traffic conditions" has a higher priority than "play music"; in a bedroom scene, "sleep aid related instructions" has a higher priority than "entertainment instructions").

[0077] The user profile dynamic update module includes:

[0078] The user profile unit is used to create a user profile based on static and dynamic data. The static data includes the user's age, gender, and basic preferences (such as elderly people's preference for "loud volume and large font"). The dynamic data includes the user's historical interaction records (such as "issued sleep-related instructions at 20:00" and "preferred soft piano music" in the past week).

[0079] The update logic unit is used to update the current instruction parameters and user feedback (such as "whether to adjust the volume") to the profile after each interaction, and optimize preference recommendations through collaborative filtering algorithms (such as automatically setting "low volume" to 15 decibels if an elderly person repeatedly adjusts the volume to 15 decibels).

[0080] The multi-device collaborative conflict handling module includes:

[0081] The conflict detection unit first identifies the set of devices and operating parameters involved in the command, and then determines the conflict through the device association matrix. For example, when the command "turn on living room music" triggers responses from both the living room speaker and the bedroom speaker, the matrix marks "multi-device response conflict". Then, it performs logical verification on the device parameter conflicts. For example, when the air conditioner turns on "26 degrees + dehumidification mode", the humidifier receives the command "turn on humidification", and the system automatically identifies "dehumidification and humidification parameters are mutually exclusive".

[0082] The conflict resolution unit prioritizes the device closest to the user's current location when dealing with multi-device response conflicts (e.g., if the user is in the living room, only the living room speaker will be triggered). If the location cannot be determined, the device with the highest historical usage frequency will be selected based on the user profile. For parameter mutual exclusion conflicts, the system makes decisions based on scenario priority. For example, in the bedroom "sleeping scenario," if the air conditioner's dehumidification need (to improve sleep comfort) is higher than the humidifier's humidification need, the system will automatically turn off the humidifier and provide feedback such as "The current air conditioner dehumidification mode is on. The humidifier has been turned off for you. Please confirm if you wish to turn it on again."

[0083] Setting up a multi-device collaborative conflict resolution module has the following effects:

[0084] 1. Avoid interactive chaos caused by multiple devices "repeated responses" and "parameter inconsistencies". For example, the 20% problem of "multiple speakers playing at the same time" in the original solution has been completely solved.

[0085] 2. Reduces the number of manual operation steps for users to coordinate devices, especially in multi-device linkage scenarios (such as smart home whole-house control), improving user operation efficiency by 40%.

[0086] The edge-cloud collaborative processing module includes:

[0087] Edge units (such as home gateways and vehicle central control systems) are used to process simple device control commands (such as "adjust the air conditioning temperature" and "turn off the lights"), with a response latency controlled within 0.3 seconds.

[0088] The cloud unit is used to handle complex tasks that require big data support (such as "searching for gas stations" and "matching a sleep aid music library"), improving efficiency through distributed computing;

[0089] The dynamic allocation unit is used to adjust the processing nodes based on network status (such as 5G / 4G / Wi-Fi). When the network is weak, high-frequency instructions (such as "play music in history") are cached at the edge to reduce cloud reliance.

[0090] The intelligent decision-making and feedback module includes:

[0091] The decision generation unit is used to send precise control commands to associated devices (such as "set the bedroom air conditioner to 26 degrees" and "turn off the living room lights").

[0092] The voice feedback unit is used to announce the execution results to the user through TTS speech synthesis technology (such as "The air conditioning adjustment and light turning off operations have been completed for you, and sleep-aid music is being played").

[0093] The anomaly handling unit is used to provide real-time feedback of anomaly information (such as "Bedroom air conditioner is not responding, please check the device status") when the device is unresponsive (e.g., air conditioner malfunction).

[0094] In practical use, the specific steps are as follows:

[0095] S1: First, the voice acquisition unit uses a noise-canceling microphone array to collect the voiceprint and intonation features of the user's voice commands; then, the scene data acquisition unit uses a temperature sensor, position sensor, and time module to obtain scene parameters; then, the device status acquisition unit establishes a communication connection with the associated device to obtain the current operating status of the device in real time.

[0096] S2: First, the data cleaning unit uses a speech recognition error correction algorithm to correct the recognition error in the speech command; then, the feature extraction unit extracts the command keywords, user voiceprint features, and time-scene correlation features from the speech data;

[0097] S3: First, the operation object and parameters of a single instruction are parsed through the basic semantic unit; then, the instruction order is identified by the temporal logic algorithm through the context semantic unit, and the dependencies between multiple instructions are associated; finally, the parsing priority is adjusted by combining the scene semantic unit with scene data.

[0098] S4: First, the user profile is composed of static and dynamic data by the profile composition unit; then, the update logic unit updates the current command parameters and user feedback to the profile after each interaction, and optimizes preference recommendations through collaborative filtering algorithm.

[0099] S5: First, the conflict detection unit identifies the set of devices and operation parameters involved in the instruction, and determines the conflict through the device association matrix; then, the device parameter conflict is logically verified; then, the conflict resolution unit responds to the conflict of multiple devices, and prioritizes the device closest to the user's current location. If the location cannot be determined, the device with the highest historical usage frequency is selected based on the user profile; for parameter mutual exclusion conflicts, the decision is made in combination with the scenario priority.

[0100] S6: The processing nodes are adjusted based on the network status through the dynamic allocation unit. When the network is weak, high-frequency instructions are cached at the edge to reduce reliance on the cloud.

[0101] S7: First, the decision generation unit sends precise control commands to the associated devices; then, the voice feedback unit uses TTS speech synthesis technology to broadcast the execution results to the user; and finally, the exception handling unit provides real-time feedback of exception information when the device does not respond.

[0102] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. An artificial intelligence voice interaction system based on big data, characterized in that, include: The multi-dimensional data acquisition module is used to collect user voice data, real-time scene data, and device status data, providing basic data support for subsequent processing; The big data preprocessing module is used to clean and extract features from the collected raw data, remove invalid data, and improve the efficiency of subsequent parsing. The multi-level semantic parsing module is used to associate the order of multiple instructions through temporal logic algorithms and adjust the parsing priority based on scenario data; The user profile dynamic update module is used to update user preferences in real time through collaborative filtering algorithms to adapt to personalized interaction needs. The multi-device collaboration conflict handling module is used to resolve operational conflicts and mutual exclusion conflicts of device parameters when multiple devices respond to the same command at the same time. The edge-cloud collaborative processing module is used to dynamically allocate processing nodes according to network status, with simple instructions processed at the edge and complex big data tasks processed in the cloud; The intelligent decision-making and feedback module is used to generate device control commands based on semantic parsing results and user profiles, and to provide feedback on the execution results to the user.

2. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The multi-dimensional data acquisition module includes: The voice acquisition unit is used to acquire the voiceprint features and intonation features of the user's voice commands using a noise-canceling microphone array; The scene data acquisition unit is used to acquire scene parameters through temperature sensors, position sensors, and a time module. The device status acquisition unit is used to establish a communication connection with associated devices and obtain the current operating status of the devices in real time.

3. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The big data preprocessing module includes: The data cleaning unit is used to correct recognition errors in voice commands through a speech recognition error correction algorithm. The feature extraction unit is used to extract instruction keywords and user voiceprint features from speech data, and time-scene correlation features from scene data.

4. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The multi-level semantic parsing module includes: Basic semantic units are used to parse the operands and parameters of a single instruction. Context semantic unit, used to identify instruction order and associate dependencies between multiple instructions through timing logic algorithms; Scene semantic unit, used to adjust parsing priority based on scene data.

5. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The user profile dynamic update module includes: A user profile building unit is used to create a user profile based on static and dynamic data. The update logic unit is used to synchronize the current instruction parameters and user feedback to the profile after each interaction, and optimize preference recommendations through collaborative filtering algorithms.

6. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The multi-device collaborative conflict handling module includes: The conflict detection unit is used to first identify the set of devices and operating parameters involved in the instruction, and then determine the conflict through the device association matrix; then it performs logical verification on the device parameter conflict. The conflict resolution unit prioritizes the device closest to the user's current location when dealing with multi-device response conflicts. If the location cannot be determined, it selects the device with the highest historical usage frequency based on the user profile. For parameter mutual exclusion conflicts, it makes decisions based on scenario priority.

7. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The edge-cloud collaborative processing module includes: Edge units are used to process simple device control commands; The cloud unit is used to process complex tasks that require big data support; The dynamic allocation unit is used to adjust the processing nodes based on the network status. When the network is weak, high-frequency instructions are cached at the edge to reduce reliance on the cloud.

8. The artificial intelligence voice interaction system based on big data according to claim 1, characterized in that, The intelligent decision-making and feedback module includes: The decision generation unit is used to send precise control commands to associated devices; The voice feedback unit is used to broadcast the execution results to the user through TTS speech synthesis technology; The exception handling unit is used to provide real-time feedback of exception information when the device is unresponsive.

Citation Information

Cited By

  • Self-adaptive voice semantic communication method based on hierarchical time sequence importance

    CN121619609A