Scene linkage sound box control system and method based on artificial intelligence

By using an AI-based scene-linked speaker control system, which leverages natural language processing and fuzzy rough set theory algorithms, the system achieves adaptive adjustment of smart devices, addressing the shortcomings of overall linkage control in smart home systems and improving user experience, system adaptability, and reliability.

CN120954399APending Publication Date: 2025-11-14GANZHOU DEHUIDA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511104786.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing smart home systems lack the ability to coordinate and control different home scenarios as a whole, resulting in cumbersome operation and difficulty in meeting the diverse and personalized needs of users in their daily lives.

Method used

An AI-based scene-linked speaker control system is adopted. Through data acquisition, preprocessing, natural language processing, and fuzzy rough set theory algorithms, it analyzes user voice commands and environmental data to generate device control commands and adaptively adjust the smart device.

Benefits of technology

It enables precise linkage control of smart devices, improves the user experience of smart homes, adapts to complex contexts and dynamic environments, reduces misunderstandings caused by semantic ambiguity, and enhances the robustness and fault tolerance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954399A_ABST
    Figure CN120954399A_ABST
Patent Text Reader

Abstract

The invention provides a scene linkage loudspeaker box control system and method based on artificial intelligence, and relates to the technical field of loudspeaker box control. Comprising a data acquisition module for acquiring a user voice instruction and environment data; the data preprocessing module is used for preprocessing the voice instruction and the environment data to obtain preprocessed context data; the artificial intelligence processing module analyzes the preprocessed context data by using a natural language processing algorithm, the sound box obtains an equipment control instruction by using a fuzzy rough set theory algorithm according to the analyzed context data, and the indoor intelligent equipment is adaptively adjusted according to the equipment control instruction; and the instruction execution module adjusts the intelligent equipment according to the equipment control instruction. According to the invention, the indoor intelligent equipment is adaptively adjusted through the natural language processing algorithm and the fuzzy rough set theory algorithm, so that the sound box can reasonably and adaptively adjust the intelligent equipment in various complex real scenes, and the overall applicability and flexibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speaker control, and in particular to a scene linkage speaker control system based on artificial intelligence and a scene linkage speaker control method based on artificial intelligence. Background Art

[0002] With the continuous popularization of the concept of smart home, more and more families begin to equip various smart devices, such as smart lighting systems, smart curtains, and smart home appliances. At present, the smart home control methods on the market often have certain limitations. Many need to be operated through a dedicated mobile phone APP. When using it, the user needs to manually open the APP and find the corresponding device control interface, and the operation is relatively cumbersome and not convenient enough. Although some smart speakers also have certain voice control functions, most of them can only achieve simple single-device command control and lack the ability to perform overall linkage adaptive control for different home scenes, making it difficult to meet the diverse and personalized life scene needs of users, such as the need to automatically adjust the states of a series of related devices according to different time periods and different activity scenes. Therefore, it is necessary to develop a more intelligent, efficient and scene-linkage-enabled speaker control system to improve the overall usage experience of smart home. Summary of the Invention

[0003] The present invention provides a scene linkage speaker control system and method based on artificial intelligence to solve the problem in the prior art of lacking the ability to perform overall linkage control for different home scenes.

[0004] On the one hand, the present invention provides a scene linkage speaker control system based on artificial intelligence, including: A data acquisition module for collecting user voice commands and environmental data.

[0005] A data preprocessing module for preprocessing the voice commands and environmental data to obtain preprocessed context data.

[0006] An artificial intelligence processing module for analyzing the preprocessed context data by using natural language processing algorithms. The speaker uses the fuzzy rough set theory algorithm based on the analyzed context data to obtain device control commands, and adaptively adjusts the smart devices in the room according to the device control commands.

[0007] An instruction execution module for adjusting the smart devices according to the device control commands.

[0008] According to the scene linkage speaker control system based on artificial intelligence provided by the present invention, the data acquisition module includes: A voice command unit for collecting user voice commands, voice error prompts, voice feedback confirmation information, voice guidance information, voice prompt information, audio playback content, and voice status announcements.

[0009] The environmental data unit is used to collect data on light intensity, sound decibels, temperature and humidity, air quality, human presence, spatial location, and time.

[0010] According to the present invention, a scene-linked speaker control system based on artificial intelligence includes a data preprocessing module comprising: The signal denoising unit is used to remove environmental noise and preserve user voice by using a bandpass filter.

[0011] The feature extraction unit is used to extract the spectral envelope features in speech and convert the speech signal into a feature vector.

[0012] The speech recognition unit is used to convert the user's voice commands into text information.

[0013] According to the present invention, an artificial intelligence-based scene-linked speaker control system includes an artificial intelligence processing module comprising: The context feature embedding unit is used to convert each word in the preprocessed context data into a vector form to obtain the context feature vector.

[0014] The part-of-speech tagging and named entity recognition unit is used to perform part-of-speech tagging and named entity recognition on context feature vectors.

[0015] The intent classification unit is used to determine the intent of preprocessed contextual data using a convolutional neural network.

[0016] The semantic analysis unit is used for context feature vector analysis to analyze the semantic structure and word-to-word dependencies of preprocessed context data.

[0017] According to the present invention, an artificial intelligence-based scene-linked speaker control system includes an artificial intelligence processing module comprising: The data discretization unit is used to discretize the preprocessed context data to obtain discrete context data.

[0018] The fuzzification processing unit is used to fuzzify discrete contextual data to obtain fuzzy membership degrees.

[0019] The fuzzy similarity calculation unit is used to calculate the fuzzy similarity of intelligent device adjustment combinations corresponding to contextual data in different scenarios based on fuzzy membership degree.

[0020] The fuzzy rough set calculation unit is used to calculate the upper and lower approximations of the fuzzy rough set based on the fuzzy similarity relationship to obtain the fuzzy rough set.

[0021] The control command generation unit is used to generate corresponding equipment control commands based on fuzzy membership degrees and according to preset decision rules, and send them to the equipment for execution and adjustment.

[0022] According to the artificial intelligence-based scene-linked speaker control system provided by the present invention, the fuzzy similarity calculation formula of the fuzzy similarity calculation unit is expressed as follows: ; In the formula, C is the attribute set, μ Fk Let x be the membership function. i and x j Let 'a' be the combination of smart device adjustments corresponding to contextual data in two scenarios, where 'a' is any attribute.

[0023] According to the artificial intelligence-based scene-linked speaker control system provided by the present invention, the fuzzy rough set calculation formula of the fuzzy rough set calculation unit is expressed as follows: Approximation under fuzzy rough sets R The membership function of (X) is: ; In the formula, I is an implication operator. It is the combination of intelligent device adjustments corresponding to contextual data in different scenarios. j The membership degree of a set X.

[0024] Approximation on fuzzy rough sets The membership function of (X) is: ; In the formula, T is a triangular modulus operator, T(a,b)=min(a,b)), and sup represents the supremum.

[0025] According to the present invention, a scene-linked speaker control system based on artificial intelligence includes an instruction execution module comprising: The instruction signal recognition unit is used to identify the identification information of the smart device in the instruction.

[0026] The signal conversion unit is used to encapsulate and package the identification information according to the protocol format of the smart device, and convert it into a signal form that the smart device can recognize and execute.

[0027] The signal feedback unit is used to wait for feedback information from the smart device and confirm whether the device has received and executed the instruction.

[0028] The scene-linked speaker control method based on artificial intelligence provided by the present invention further includes a device communication module: The instruction extraction unit is used to extract key target devices from device control instructions.

[0029] The Wi-Fi protocol unit is used to establish a link for Wi-Fi communication with the target device.

[0030] The Zigbee protocol unit is used to integrate with the Zigbee network where the target device is located to establish a communication connection.

[0031] The Bluetooth protocol unit is used to identify the Bluetooth signal of the target device and establish a Bluetooth connection.

[0032] The Z-Wave protocol unit is used to transmit data to security monitoring equipment in real time and trigger alarms.

[0033] On the other hand, the present invention also provides a scene-linked speaker control method based on artificial intelligence, including: S1: Collect user voice commands and environmental data.

[0034] S2: Preprocess the voice commands and environmental data to obtain preprocessed contextual data.

[0035] S3: Used to analyze the preprocessed context data using natural language processing algorithms. The speaker uses fuzzy rough set theory algorithm to obtain device control commands based on the analyzed context data, and adaptively adjusts the smart devices in the room according to the device control commands.

[0036] S4: Adjust the smart device according to the device control command.

[0037] This invention provides an artificial intelligence-based scene-linked speaker control system and system. It analyzes pre-processed contextual data using natural language processing algorithms, and the speaker adaptively adjusts indoor smart devices based on the analyzed contextual data using fuzzy rough set theory algorithms. The beneficial effects achieved are as follows: Natural language processing algorithms excel at parsing the grammatical structure of text, recognizing intent, and extracting key entities. They can perform preliminary analysis of the contextual data in the form of natural language input by users. By analyzing information systems built from a large amount of historical data and combining operations such as fuzzification and attribute reduction, they can uncover the actual adjustment range and strategies of smart devices under different fuzzy semantic expressions. The combination of these two approaches can more accurately grasp user intent, avoid misunderstandings caused by semantic fuzziness, and allow speakers to more accurately know the indoor smart device adjustment status that users expect.

[0038] In real life, contextual data is often highly complex. Natural language processing (NLP) algorithms can integrate and analyze these diverse textual descriptions to identify key elements. Fuzzy rough set theory algorithms can handle the fuzzy relationships and uncertain connections between these complex contextual elements. For example, saying "brighter" might have different effects on adjusting the brightness of lights in different seasons and environments. This algorithm can adapt to the current context parsed by NLP based on past data reflecting different contextual adjustment patterns. This allows the speaker to reasonably and adaptively adjust the smart device in various complex real-world scenarios, improving its overall applicability and flexibility.

[0039] The natural language processing algorithm of this invention can continuously collect and analyze new contextual text data, and promptly respond to new user needs and changes in usage scenarios. The fuzzy rough set theory algorithm can continuously update its information system based on new data, reassess the importance of attributes, and generate new decision rules. The combination of the two enables the speaker's strategy for adjusting indoor smart devices to keep pace with the times, continuously optimize, and always remain in an optimal state that conforms to the current actual situation, better adapting to the dynamically changing home environment and user needs.

[0040] When faced with noisy contextual data, natural language processing algorithms strive to restore the correct semantics through error correction and completion. Fuzzy rough set theory algorithms themselves have good tolerance for the fuzziness and certain degree of imprecision of data. When the two are combined, the system can still reliably analyze contextual data and make reasonable intelligent device adjustment decisions even with certain data quality issues. This reduces the risk of adjustment errors caused by individual data anomalies or imperfect expressions, and enhances the robustness and fault tolerance of the entire system. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0042] Figure 1 This is a schematic diagram of a scene-linked speaker control system based on artificial intelligence provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a scene-linked speaker control method based on artificial intelligence, provided in an embodiment of the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0044] The following is combined with Figures 1-2 This invention describes an artificial intelligence-based scene-linked speaker control system and system.

[0045] Figure 1 This is a schematic diagram of a scene-linked speaker control system based on artificial intelligence, provided in an embodiment of the present invention.

[0046] like Figure 1 As shown in the figure, an embodiment of the present invention provides a scene-linked speaker control system based on artificial intelligence. The system includes: The data acquisition module is used to collect user voice commands and environmental data.

[0047] As the interface for user interaction, the acquisition module has a built-in high-sensitivity microphone and a high-performance speaker for providing feedback on system prompts and playing audio content.

[0048] The environmental data unit includes: light intensity data, sound decibel data, temperature and humidity data, air quality data, human presence data, spatial location data, and time data.

[0049] Environmental data is processed using light sensors to detect ambient light intensity and determine whether it is day or night; sound sensors to monitor ambient noise levels to help determine the quietness or noise level of the scene; infrared human sensors to detect the presence of people; and temperature and humidity sensors to determine the comfort level of the indoor environment. The system also receives status information from other devices in the smart home, such as the on / off status of smart door locks and the opening / closing status of smart curtains. This multi-source data is used to perceive and make a preliminary judgment about the current scene.

[0050] The data preprocessing module is used to preprocess voice commands and environmental data.

[0051] The signal denoising unit removes ambient noise using a bandpass filter, preserves the frequency range of human speech, determines the start and end points of speech, and avoids processing invalid silence segments. It also adjusts the signal amplitude to eliminate the effects caused by differences in speaker volume.

[0052] The feature extraction unit is used to simulate the differences in human ear perception of different frequencies using Mel frequency cepstral coefficients, extract the spectral envelope features of speech, and convert the preprocessed speech signal into a feature vector.

[0053] The speech recognition unit is used to accurately convert users' voice commands into text information. The core task of the speech recognition module is to convert the acoustic features of human speech into computer-understandable text information. The speech recognition unit includes a signal processing subunit, a pattern recognition subunit, and a natural language processing subunit.

[0054] The signal processing subunit is used to convert continuous speech waveforms into discrete acoustic feature vectors.

[0055] The pattern recognition subunit is used to compare acoustic features with a pre-trained speech model to identify the most likely phonemes or syllables.

[0056] The speech enhancement subunit is used to determine the direction of the speech signal through multiple microphone arrays using beamforming technology, enhance the target speech, and suppress noise from other directions.

[0057] The intelligent device adjustment combination library subunit, which corresponds to contextual data in different speech scenarios, is used to collect representative dialects from various regions of China and the pronunciation characteristics of languages ​​from different countries and ethnic groups, and to regularly update the intelligent device adjustment combination library corresponding to contextual data in different scenarios.

[0058] The artificial intelligence processing module is used to analyze the pre-processed contextual data using natural language processing algorithms. Based on the contextual data, the speaker uses fuzzy rough set theory algorithms to adaptively adjust the smart devices in the room and obtain device control commands.

[0059] Natural language processing algorithms can identify the intent in instructions and determine whether to execute single device control or scene-linked control.

[0060] Natural language processing algorithms analyze preprocessed contextual data, including: Set the preprocessed context data as a text sequence T=[ω1, ω2, ..., ω]. n ], where ω i Let fi represent the i-th word in the text, and n be the text length. The contextual feature vector F = [f1, f2, ..., fn]. k] f j Let j represent the j-th contextual feature, such as the current time, ambient light intensity, or whether there are people indoors, which are represented in numerical or discrete category form, and k is the number of contextual features.

[0061] Word embedding is used to convert each word in the text into a vector form. Simultaneously, consideration is given to how to integrate contextual feature information to enable the computer to more comprehensively understand the semantic relationships between words and their context. Assuming the word vector dimension is d, for a word ω... i Its corresponding word vector v ωi ∈R d The word vector sequence of the entire text T can be represented as V. T =[v ω1 v ω2 , ..., v ωn ].

[0062] For the context feature vector F, a simple concatenation method is used to fuse it with the word vectors, and each word vector v is combined. ωi The fusion vector is concatenated with the contextual feature vector F to obtain a new fusion vector. ∈R d+k The fused vector sequence of the entire concatenated text T is represented as follows: =[ , ..., ].

[0063] Taking the Skip-gram model in Word2Vec as an example, its goal is to predict the context words based on the center word. The training objective function can be expressed as: ; In the formula, D represents all word pairs in the corpus. Given a center word w and contextual features F, the probability of predicting the context word c is calculated. In a specific implementation, this involves fusing operations such as the inner product of word vectors and normalizing using the softmax function. It can be represented as: ; in, Y is the feature function, which is the set of all possible label sequences. The most likely part-of-speech or named entity label sequence is determined by maximizing this probability. At this time, the judgment will be affected by context features.

[0064] Convolutional neural networks determine the intent of text, considering the influence of contextual features on intent judgment. This is applied to word vector sequences incorporating contextual features. Convolution operations are performed using convolution kernels of different sizes k, with the kernel size set to m×(d+k), where m is the number of words covered by the kernel (incorporating contextual features, the dimension becomes d+k). The convolution calculation formula is as follows: ; in, It is the result of convolving the i-th position with the k-th convolution kernel. It is the l-th weight of the k-th convolutional kernel. It is the element at the corresponding position in the fused vector sequence, b k It is the bias term, and ReLU is the activation function.

[0065] After pooling and other operations, the resulting feature vectors are input into a fully connected layer, where they are classified using the softmax function. The softmax function is calculated as follows: ; Where p(y i ) is the probability that the intent belongs to the i-th class, zi is the corresponding score of the fully connected layer output, and C is the total number of intent categories. The integration of contextual features can help to judge the intent more accurately. For example, if someone says "make it darker", the intent classification result will be different depending on whether the current scenario is watching a movie or sleeping. In the former case, it may be to dim the lights to create a movie atmosphere, while in the latter case, it may be to dim the lights to help with sleep.

[0066] The semantic analysis and dependency parsing unit utilizes methods based on transition systems or graph neural networks, combined with grammatical rules and contextual features, to analyze the semantic structure of text and the dependency relationships between words. For example, when constructing a parsing tree, in addition to relying on the grammatical structure rules of the text itself, contextual features are also considered. In the current context of a "family gathering," the phrase "turn the music up a bit" can be more accurately understood to refer to the music currently playing to create a party atmosphere, rather than other unrelated music. Although there is no simple, universal formula for this, contextual features can assist in a deeper and more accurate analysis of semantics, clarifying the subject, verb, and object grammatical components of the sentence, as well as the modifying and governing relationships between words.

[0067] The contextual reasoning and knowledge fusion unit combines external knowledge graphs (if available) with predefined domain knowledge rules, while making full use of contextual features to fuse and reason about the analyzed intent, entities, and other information, in order to better understand the user's complete needs in the current context.

[0068] In a specific embodiment, given that the ambient light is dim, it is nighttime, and the user says "play some music," the system combines knowledge from the knowledge graph about suitable music types for nighttime and music styles often associated with dim environments to recommend playing soothing and quiet music. Through this reasoning and fusion, the system infers the information implied in the text and its association with specific smart devices and application scenarios, thus better meeting the user's needs in the current context.

[0069] The device control instruction generation unit is used to generate corresponding control instructions based on semantics, intent, and reasoning results, in accordance with the instruction format acceptable to the smart device and taking into account contextual features.

[0070] The speaker adaptively adjusts indoor smart devices using fuzzy rough set theory algorithms based on contextual data. The specific steps to obtain device control commands include: Let the information system be S = (U, A, V, f), where U is the universe of discourse, that is, the set of intelligent device adjustment combinations corresponding to contextual data in different scenarios. These intelligent device adjustment combinations corresponding to contextual data in different scenarios are instances of intelligent device adjustment in different contexts, which can be represented as U = {x1, x2, ..., x}. n} (n is the number of intelligent device adjustment combinations corresponding to contextual data in different scenarios). A is a set of attributes, consisting of conditional attribute C and decision attribute D, i.e., A=C∪D. V is a set of attribute values, with different attributes having their corresponding value ranges. f is an information function used to determine the value of each object under each attribute. For example, f(xi, a) represents the value of the intelligent device adjustment combination xi corresponding to contextual data in different scenarios under attribute a (a∈A).

[0071] The data discretization unit is used to discretize the preprocessed context data to obtain discrete context data. For attributes with continuous values ​​in conditional attribute C, discretization is required to transform them into discrete values ​​for subsequent analysis using fuzzy rough set theory. Commonly used discretization methods include equal-width discretization and equal-frequency discretization.

[0072] Taking constant-width discretization as an example, let the continuous attribute a∈C have a range of values ​​[min a max a To divide the data into k intervals, the width of each interval is... Discretized attribute value v ij Let j be the j-th interval, and let x be the combination of smart devices adjusted according to the contextual data in different scenarios. i The interval to which the attribute belongs is determined by the value f(xi, a) on attribute a. If f(xi, a) ∈ [min a +(j−1)ω,mina+jω), then v ij These are the corresponding discrete values.

[0073] The fuzzification processing unit is used to fuzzify discrete contextual data to obtain fuzzy membership degrees.

[0074] Suppose that for a certain attribute a, m fuzzy sets F1, F2, ..., Fm are defined. m For contextual data in different scenarios, the value v of the intelligent device adjustment combination xi on attribute a is determined. ij Its membership function μ belongs to the fuzzy set Fk Fk (v ij The calculation is based on the specific membership function form.

[0075] The fuzzy similarity calculation unit is used to calculate the intelligent device adjustment combination corresponding to contextual data in different scenarios based on fuzzy membership degree.

[0076] For the smart device adjustment combination x corresponding to the contextual data in the two scenarios i and x j The fuzzy similarity relation R(x) on the attribute set C i x j The calculation formula is: ; In the formula, C is the attribute set, μ Fk Let x be the membership function. i and x j Let 'a' be the intelligent device adjustment combination corresponding to contextual data in two scenarios, where 'a' is an arbitrary attribute. For each attribute 'a', first take the maximum value of the membership degree of the intelligent device adjustment combination corresponding to contextual data in two different scenarios under each fuzzy set of that attribute, and then take the minimum value of the calculated results of all attributes to obtain the fuzzy similarity value between them. This value is in the range of [0,1], and the larger the value, the more similar they are.

[0077] The fuzzy rough set computation unit is used to calculate the upper and lower approximations of fuzzy rough sets based on fuzzy similarity relations using specific implication operators and triangular modulus operators, in order to characterize the classification of decision attributes based on conditional attributes.

[0078] The upper and lower approximations of fuzzy rough sets are defined by fuzzy similarity relations. This is a core concept in fuzzy rough set theory and is used to classify and describe decision attribute D based on condition attribute C.

[0079] Let X⊆U be a subset of the universe of discourse U, and its fuzzy lower approximation R The membership function of (X) is: ; In the formula, I is an implication operator. It is the combination of intelligent device adjustments corresponding to contextual data in different scenarios. j The membership degree of a set X is 1 if xj belongs to X, and 0 otherwise. inf denotes the infimum, which means taking all xj. j The minimum value in the calculation results.

[0080] Fuzzy approximation The membership function of (X) is: ; In the formula, T is a triangular modulus operator, T(a,b)=min(a,b)), and sup represents the supremum, which is to take the maximum value among all the calculated results of xj.

[0081] The attribute importance assessment unit calculates the importance of conditional attributes to decision attributes, removes unimportant redundant attributes, and obtains a reduced attribute set, focusing on key contextual data features.

[0082] Let C′ be a subset of C, and POS C D) represents the positive domain of D relative to C, and the importance σ of attribute a∈C relative to D. CD (a) The calculation formula is: ; in R C (D) is a fuzzy approximation of decision attribute D based on attribute set C. This is a fuzzy approximation of the decision attribute D based on the attribute set C−{a} after removing attribute a.

[0083] By calculating the importance of each attribute, the attributes with higher importance are selected to form a reduced attribute set C*, and redundant attributes that have little impact on decision-making are removed.

[0084] The control command generation unit is used to acquire and fuzzify the current context data, and generate corresponding device control commands according to the matching decision rules, the adjustment range and protocol format of the smart device, and send them to the device for execution and adjustment.

[0085] The device communication module is used to establish communication connections with smart devices based on device control commands. The communication protocols used in speaker control include Wi-Fi, Zigbee, Bluetooth, and Z-Wave. The device communication module must fully understand the protocol types corresponding to different devices in order to effectively perform communication connection operations.

[0086] The instruction extraction unit processes incoming device control instructions, extracting key target device information such as the device's specific model, location, or unique device identifier. This step is crucial because only by identifying the specific smart device to communicate with can subsequent precise connection procedures be carried out. For example, this information is used to distinguish whether to connect to a smart speaker in the living room or a smart air conditioner in the bedroom.

[0087] The Wi-Fi protocol unit is used to search for surrounding Wi-Fi networks using the speaker's built-in Wi-Fi communication function. It locates the network the target device is connected to and performs network authentication, handshake, and other operations according to Wi-Fi communication standards, establishing a stable and reliable link between the devices. Wi-Fi offers advantages such as relatively wide coverage and high transmission speeds, making it popular among many smart devices. The 2.4GHz band signal has strong penetration, easily crossing indoor walls and obstacles, enabling large-area coverage in ordinary homes and offices, ensuring that smart speakers, smart cameras, and other devices in different areas can access the network. Wi-Fi 6 and later standards theoretically offer transmission speeds of several Gbps, greatly ensuring the smoothness of smart speakers accessing online audio resources and the timeliness of smart cameras transmitting high-definition video, facilitating remote control and viewing by users. This allows smart devices to be efficiently integrated into smart home systems, providing users with convenient smart services.

[0088] The Zigbee protocol unit is used to send network request signals conforming to the Zigbee protocol specification, such as requesting to form a network or join a network, through the speaker's built-in Zigbee module. This allows the unit to merge with the Zigbee network of the target device, thereby establishing a communication connection. During this process, many details, such as network address allocation and channel selection, must be handled to ensure smooth communication. The Zigbee protocol boasts advantages such as low power consumption and flexible networking. Due to its unique design, the device consumes very little power during operation, which is significant for smart devices that rely on batteries or have stringent energy consumption requirements. For example, wireless bulbs in smart lighting systems and various smart sensors can operate stably for extended periods using button batteries, reducing user maintenance costs and frequency. It supports various topologies such as star, tree, and mesh, allowing for flexible networking based on the actual layout of home devices. This facilitates communication and collaborative control between devices, and the addition of new devices is also simple and easy. Therefore, it is widely used for short-range communication between smart lighting systems, smart sensors, and other devices, helping to create a convenient and comfortable smart home environment.

[0089] The Bluetooth protocol unit, or device communication module, enables Bluetooth, scans for nearby Bluetooth devices, identifies the target device's Bluetooth signal, and initiates a pairing request. After confirmation from the device and verification of the necessary pairing code, a Bluetooth connection is successfully established, enabling data exchange between the two devices. The Bluetooth protocol is primarily used for short-range communication, typically operating at 2.4GHz, and can establish a stable and efficient connection within a range of approximately 10 meters.

[0090] The Z-Wave protocol unit boasts advantages in security and stability. In terms of security, it employs advanced encryption technology to strongly encrypt sensitive data such as smart lock unlock codes and security monitoring footage. Only authorized recipients with the decryption key can process the data, effectively preventing malicious external attacks and protecting home privacy and security. The Z-Wave protocol operates on a relatively independent frequency band, avoiding common interference sources and maintaining a stable communication link even in complex wireless environments. This is crucial for real-time data transmission from security monitoring devices and for inter-device alarm linkage, ensuring reliable operation of home security features in critical moments.

[0091] The continuous monitoring unit constantly monitors the status of the communication link. In real-world scenarios, network interference, equipment malfunctions, or other unexpected events may cause communication link interruptions or signal weakening. Once a communication anomaly is detected, the module must promptly take appropriate measures to attempt to restore the connection, such as resending the connection request, switching to a backup communication channel, or prompting the user to check the device's network connectivity. This ensures that communication between the entire system and the smart device remains normal and available, allowing subsequent operations based on device control commands to be accurately transmitted to and executed by the smart device.

[0092] Furthermore, as smart devices in the home are constantly being updated or new devices are added, the device communication module must also have good compatibility and scalability, be able to quickly adapt to new devices and new communication protocols, and continuously optimize its own communication connection methods and processing mechanisms to ensure that the entire smart home scene linkage system always maintains an efficient and stable operating state in terms of device communication.

[0093] The instruction execution module is used to adjust the smart device according to the device control instructions.

[0094] The command signal recognition unit parses received device control commands and identifies the specific smart device involved in the command. For example, it uses a unique device code or name to determine whether it's a smart speaker in the living room, a smart lighting system in the bedroom, or a smart appliance in the kitchen. Simultaneously, it parses the specific actions to be performed on the device. For a speaker, actions might include play, pause, switch tracks, adjust volume, or switch sound effects modes. For a smart lighting system, actions might include turning lights on, turning lights off, adjusting brightness, or changing light color. For other smart appliances, actions might include starting, stopping, or adjusting operating parameters.

[0095] The signal conversion unit is used to encapsulate and package the parsed operation instructions according to the corresponding protocol format, converting them into a signal form that the device can recognize and execute. If communicating with the smart speaker via the Wi-Fi protocol, the volume adjustment instructions will be encoded and encapsulated according to the Wi-Fi control instruction format specified by the speaker manufacturer, and then sent to the speaker.

[0096] The signal feedback unit waits for feedback from the smart device to confirm whether the device has accurately received the instruction and successfully executed the corresponding operation. If the device returns a successful execution feedback, the module records the execution status and feeds the relevant information back to other parts of the system, such as informing the AI ​​processing module that the instruction has been executed and subsequent interactions or scene status updates can proceed. If the device reports an execution failure, the module analyzes possible causes, such as communication failure, a problem with the device itself, or an incorrect instruction format, and attempts to resend the instruction or take appropriate remedial measures, such as switching the communication link or prompting the user to check the device status, to ensure the normal operation of the entire system and the smooth realization of scene linkage.

[0097] In complex interactive scenarios, multiple smart devices may need to perform different operations simultaneously or in a certain order. The instruction execution module needs to have good task scheduling and coordination capabilities to ensure proper coordination between the operations of each device. For example, in a home theater scenario, the smart curtains need to be closed first, then the projector needs to be turned on, and then the speakers need to be adjusted to the appropriate sound effects and volume. Throughout the process, the instruction execution module must send corresponding instructions to each device in an orderly manner, strictly controlling the order and time interval of the operations, so that the devices can work together to create an ideal viewing environment.

[0098] As the system continues to run and be used, the instruction execution module will continuously optimize its execution strategy and adaptability to different devices based on the actual execution situation and feedback data from the devices. For example, by statistically analyzing data such as instruction response time and success rate of different devices, it can dynamically adjust the frequency of instruction sending, retransmission mechanism, and communication parameters to better adapt to different usage scenarios and device status changes, thereby improving the accuracy and reliability of the entire system's adjustment of smart devices.

[0099] In summary, this embodiment provides an AI-based scene-linked speaker control system. It analyzes pre-processed contextual data using natural language processing algorithms, and the speaker adaptively adjusts indoor smart devices based on the analyzed contextual data using fuzzy rough set theory algorithms. The beneficial effects achieved are as follows: Natural language processing algorithms excel at parsing the grammatical structure of text, recognizing intent, and extracting key entities. They can perform preliminary analysis of the contextual data in the form of natural language input by users. By analyzing information systems built from a large amount of historical data and combining operations such as fuzzification and attribute reduction, they can uncover the actual adjustment range and strategies of smart devices under different fuzzy semantic expressions. The combination of these two approaches can more accurately grasp user intent, avoid misunderstandings caused by semantic fuzziness, and allow speakers to more accurately know the indoor smart device adjustment status that users expect.

[0100] Based on the same general inventive concept, this invention also protects a big data-based reverse express delivery status identification and tracking management system. The big data-based reverse express delivery status identification and tracking management system provided by this invention will be described below. The big data-based reverse express delivery status identification and tracking management system described below can be referred to in correspondence with the artificial intelligence-based scene linkage speaker control system described above.

[0101] Figure 2 This is a flowchart illustrating a scene-linked speaker control method based on artificial intelligence, provided in an embodiment of the present invention.

[0102] like Figure 2 As shown, an artificial intelligence-based scene-linked speaker control method includes: S1: Collect user voice commands and environmental data.

[0103] S2: Preprocess the voice commands and environmental data to obtain preprocessed contextual data.

[0104] S3: Used to analyze the preprocessed context data using natural language processing algorithms. The speaker uses fuzzy rough set theory algorithm to obtain device control commands based on the analyzed context data, and adaptively adjusts the smart devices in the room according to the device control commands.

[0105] S4: Adjust the smart device according to the device control command.

[0106] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A scene-linked speaker control system based on artificial intelligence, characterized in that, include: The data acquisition module is used to collect user voice commands and environmental data; The data preprocessing module is used to preprocess voice commands and environmental data to obtain preprocessed contextual data; The artificial intelligence processing module is used to analyze the preprocessed context data using natural language processing algorithms. The speaker uses fuzzy rough set theory algorithm to obtain device control commands based on the analyzed context data, and adaptively adjusts the smart devices in the room according to the device control commands. The instruction execution module is used to adjust the smart device according to the device control instructions.

2. The scene-linked speaker control system based on artificial intelligence according to claim 1, characterized in that, The data acquisition module includes: The voice command unit is used to collect user voice commands, voice error prompts, voice feedback confirmation information, voice guidance information, voice prompt information, audio playback content, and voice status broadcasts; The environmental data unit is used to collect data on light intensity, sound decibels, temperature and humidity, air quality, human presence, spatial location, and time.

3. The scene-linked speaker control system based on artificial intelligence according to claim 1, characterized in that, The data preprocessing module includes: The signal denoising unit is used to remove environmental noise and preserve user voice by using a bandpass filter; The feature extraction unit is used to extract the spectral envelope features in speech and convert the speech signal into a feature vector. The speech recognition unit is used to convert the user's voice commands into text information.

4. The scene-linked speaker control system based on artificial intelligence according to claim 1, characterized in that, The artificial intelligence processing module includes: The context feature embedding unit is used to convert each word in the preprocessed context data into a vector form to obtain a context feature vector. The part-of-speech tagging and named entity recognition unit is used to perform part-of-speech tagging and named entity recognition on the context feature vector; An intent classification unit is used to determine the intent of the preprocessed contextual data using a convolutional neural network. The semantic analysis unit is used to analyze the semantic structure and word-to-word dependencies of the preprocessed contextual data using the context feature vector.

5. The scene-linked speaker control system based on artificial intelligence according to claim 1, characterized in that, The artificial intelligence processing module includes: The data discretization unit is used to discretize the preprocessed context data to obtain discrete context data. The fuzzification processing unit is used to fuzzify the discrete context data to obtain fuzzy membership degrees; The fuzzy similarity calculation unit is used to calculate the fuzzy similarity of intelligent device adjustment combinations corresponding to contextual data in different scenarios based on the fuzzy membership degree. The fuzzy rough set calculation unit is used to calculate the upper and lower approximations of the fuzzy rough set based on the fuzzy similarity relationship to obtain the fuzzy rough set. The control command generation unit is used to generate corresponding equipment control commands based on the fuzzy membership degree and according to preset decision rules, and send them to the equipment for execution and adjustment.

6. The scene-linked speaker control system based on artificial intelligence according to claim 5, characterized in that, The fuzzy similarity calculation formula of the fuzzy similarity calculation unit is expressed as follows: ; In the formula, C is the attribute set, and μ Fk Let x be the membership function. i and x j Let 'a' be the combination of smart device adjustments corresponding to contextual data in two scenarios, where 'a' is any attribute.

7. A scene-linked speaker control system based on artificial intelligence according to claim 5, characterized in that, The fuzzy rough set calculation formula of the fuzzy rough set calculation unit is expressed as follows: Approximation under fuzzy rough sets R The membership function of (X) is: ; In the formula, I is an implication operator. It is the combination of intelligent device adjustments corresponding to contextual data in different scenarios. j The degree of membership to set X; Approximation on fuzzy rough sets The membership function of (X) is: ; In the formula, T is a triangular modulus operator, T(a,b)=min(a,b)), and sup represents the supremum.

8. The scene-linked speaker control system based on artificial intelligence according to claim 1, characterized in that, The instruction execution module includes: The command signal recognition unit is used to identify the identification information of the smart device in the command; The signal conversion unit is used to encapsulate and package the identification information according to the protocol format of the smart device, and convert it into a signal form that the smart device can recognize and execute; The signal feedback unit is used to wait for feedback information from the smart device and confirm whether the device has received and executed the instruction.

9. A scene-linked speaker control system based on artificial intelligence according to claim 1, characterized in that, It also includes a device communication module: The instruction extraction unit is used to extract key target devices from the device control instructions; The Wi-Fi protocol unit is used to establish a link for Wi-Fi communication with the target device; The Zigbee protocol unit is used to integrate with the Zigbee network where the target device is located to establish a communication connection; The Bluetooth protocol unit is used to identify the Bluetooth signal of the target device and establish a Bluetooth connection; The Z-Wave protocol unit is used to transmit data to security monitoring equipment in real time and trigger alarms.

10. A scene-linked speaker control method based on artificial intelligence, applied to a scene-linked speaker control system based on artificial intelligence as described in any one of claims 1 to 9, characterized in that, The speaker control method includes: S1: Collects user voice commands and environmental data; S2: Preprocess the voice commands and environmental data to obtain preprocessed context data; S3: Used to analyze the preprocessed context data using natural language processing algorithms. The speaker uses fuzzy rough set theory algorithm to obtain device control commands based on the analyzed context data, and adaptively adjusts the smart devices in the room according to the device control commands. S4: Adjust the smart device according to the device control command.