Bulb lamp interaction method and device based on audio recognition and storage medium
By using audio recognition technology and pre-trained models, the system responds to user touch commands, processes audio data, generates target audio data, and controls lighting effects. This solves the problem of inconvenient control methods for smart home devices and enhances user experience and interactive fun.
Patent Information
- Application Number
- CN202510876039.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-18
AI Technical Summary
Existing smart home device control methods are insufficient to meet users' needs for convenience, flexibility, and fun, especially their diverse needs in audio and visual interaction.
By using audio recognition technology, the system responds to user touch commands to acquire audio data, performs audio embedding and fusion, generates target audio data using a pre-trained audio recognition model, and combines it with light control commands to achieve interactive effects between audio and light.
It enables users to achieve audio interaction and control in smart homes through simple touch operations, enhancing the user experience and auditory experience, and showcasing the intelligence level of smart home devices.
Smart Images

Figure CN120973285A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and in particular relates to an interactive method, device and storage medium for a bulb lamp based on audio recognition. Background Technology
[0002] With the rapid development of smart home technology, traditional lighting devices such as light bulbs are gradually transforming into intelligent devices, integrating not only lighting functions but also audio playback, Bluetooth connectivity, voice control, and other features. Audio and visual interaction are becoming an indispensable part of people's daily lives. As users' needs for smart device functions diversify, existing control methods are no longer sufficient to meet users' demands for convenience, flexibility, and enjoyment. Summary of the Invention
[0003] This application provides an interactive method, device, and storage medium for bulb lights based on audio recognition, which can solve the above-mentioned problems.
[0004] In a first aspect, embodiments of this application provide an interactive method for a light bulb based on audio recognition, including:
[0005] Responding to a user touch command, the system acquires the audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the identifier of the virtual audio component corresponding to the current bulb;
[0006] Obtain audio timing data corresponding to several segments of audio data output by the current virtual audio component;
[0007] Audio embedding processing is performed on several segments of audio data output by the current virtual audio component and the audio time sequence data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix;
[0008] The audio fusion embedding matrix is input into a pre-trained audio recognition model to obtain the first target audio data.
[0009] The audio data output by the current virtual audio component and the first target audio data are fused together to obtain the second target audio data;
[0010] Play the second target audio data.
[0011] Further, the audio embedding process performed on several segments of audio data output by the current virtual audio component and the corresponding audio temporal data of the current virtual audio component to obtain an audio fusion embedding matrix includes:
[0012] Obtain the embedding vector corresponding to the audio data and the embedding vector corresponding to the audio temporal data;
[0013] The embedding vector corresponding to the audio data is input into a preset audio feature extraction model to obtain the embedding feature vector corresponding to the audio data.
[0014] The embedded feature vector corresponding to the audio data and the embedded vector corresponding to the audio temporal data are encoded and fused to obtain the audio fusion embedding matrix.
[0015] Further, before inputting the audio fusion embedding matrix into the pre-trained audio recognition model to obtain the first target audio data, the following steps are included:
[0016] Acquire the audio training data and the audio training data;
[0017] The audio training data is segmented to obtain several audio training data slices;
[0018] Extract a first training data slice and a second training data slice from the audio training data slice; wherein the second training data slice corresponds to the first training data slice;
[0019] The first training data slice is subjected to audio embedding processing to obtain the audio fusion embedding matrix;
[0020] Based on the audio fusion embedding matrix corresponding to the first training data slice, the second training data slice, the preset loss function, and the preset optimization algorithm, the audio recognition model is iteratively trained until the recognition result of the audio fusion embedding matrix corresponding to the first training data slice meets the preset training termination condition, thereby obtaining the pre-trained audio recognition model.
[0021] Further, the step of fusing the audio data output by the current virtual audio component and the first target audio data to obtain the second target audio data includes:
[0022] The audio data output by the current virtual audio component and the first target audio data are subjected to spectrum conversion to obtain current spectrum data and first spectrum data;
[0023] Based on the current spectrum data, the first spectrum data, and the preset spectrum matching algorithm, a fusion marker corresponding to the current spectrum data is added to the first spectrum data;
[0024] The second target audio data is obtained based on the audio data output by the current virtual audio component, the first target audio data, and the first spectrum data.
[0025] Furthermore, prior to responding to the user's touch command, the process includes:
[0026] Displays a map showing the arrangement of bulbs corresponding to the virtual audio device; wherein, the virtual audio device comprises several virtual audio components;
[0027] Obtain the relative position of the bulb;
[0028] Based on the relative positions of the bulbs and the correspondence between the preset virtual audio component identifiers and their relative positions, the virtual audio component identifiers corresponding to each bulb are determined.
[0029] Send the virtual audio component identifier corresponding to each of the bulbs to the corresponding bulbs;
[0030] Furthermore, the user touch command also includes touch duration information and touch pressure information, and the acquisition of audio data output by the current virtual audio component includes:
[0031] The current virtual audio component is triggered based on the virtual audio component identifier corresponding to the current bulb light;
[0032] Based on the touch duration information and the touch pressure information, the audio data output by the current virtual audio component is obtained; wherein, the touch duration information is used to determine the duration corresponding to the audio data, and the touch pressure information is used to determine the loudness corresponding to the audio data.
[0033] Further, playing the second target audio data includes:
[0034] The second target audio data is subjected to spectrum conversion to obtain second spectrum data;
[0035] Based on the second spectrum data, a lighting control command is generated; wherein, the lighting control command includes at least bulb flashing parameters and bulb dynamic color parameters;
[0036] Send the lighting control command to the bulb.
[0037] Secondly, embodiments of this application provide an interactive bulb light device based on audio recognition, comprising:
[0038] The first processing unit is used to respond to user touch commands and obtain audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the virtual audio component identifier corresponding to the current bulb;
[0039] The acquisition unit is used to acquire audio timing data corresponding to several segments of audio data output by the current virtual audio component;
[0040] The second processing unit is used to perform audio embedding processing on several segments of audio data output by the current virtual audio component and the audio time sequence data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix.
[0041] The third processing unit is used to input the audio fusion embedding matrix into a pre-trained audio recognition model to obtain the first target audio data.
[0042] The fourth processing unit is used to fuse the audio data output by the current virtual audio component with the first target audio data to obtain the second target audio data;
[0043] The playback unit is used to play the second target audio data.
[0044] Thirdly, embodiments of this application provide an interactive bulb light device based on audio recognition, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in the first aspect above.
[0045] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect above.
[0046] In this embodiment, in response to a user touch command, the system acquires the audio data output by the current virtual audio component; acquires the corresponding audio timing data; performs audio embedding processing on the audio data and the corresponding audio timing data to obtain an audio fusion embedding matrix; obtains the first target audio data through the audio fusion embedding matrix; fuses the audio data output by the current virtual audio component and the first target audio data to obtain the second target audio data; and plays the second target audio data. This method allows users to achieve audio interaction and control in smart homes with simple touch actions. Through sophisticated audio processing technology, it ensures a rich auditory experience for users, improving their overall user experience. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a schematic flowchart of the bulb light interaction method based on audio recognition provided in the first embodiment of this application;
[0049] Figure 2 This is a schematic flowchart of steps S107 to S1010 in the bulb lamp interaction method based on audio recognition provided in the first embodiment of this application;
[0050] Figure 3 This is a schematic flowchart of the S103 method for interactive bulb lamps based on audio recognition provided in the first embodiment of this application;
[0051] Figure 4 This is a schematic flowchart of the bulb lamp interaction method S105 based on audio recognition provided in the first embodiment of this application;
[0052] Figure 5 This is a schematic flowchart of the bulb lamp interaction method S106 based on audio recognition provided in the first embodiment of this application;
[0053] Figure 6 This is a schematic diagram of the interactive bulb light device based on audio recognition provided in the second embodiment of this application;
[0054] Figure 7 This is a schematic diagram of an interactive bulb light device based on audio recognition provided in the third embodiment of the present invention. Detailed Implementation
[0055] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0056] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0057] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0058] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0059] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0060] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0061] Please see Figure 1 , Figure 1 This is a schematic flowchart of the bulb light interaction method based on audio recognition provided in the first embodiment of this application. The subject executing the bulb light interaction method based on audio recognition is a device with bulb light interaction function based on audio recognition, such as... Figure 1 The audio recognition-based bulb interaction method shown includes:
[0062] S101: Respond to the user's touch command and obtain the audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the identifier of the virtual audio component corresponding to the current bulb.
[0063] In this embodiment, the user touch command is generated when the user touches the current bulb. The user touch command is a signal generated by the user touching the bulb, which is used to instruct a specific operation.
[0064] User touch commands must include at least the identifier of the virtual audio component corresponding to the current bulb. The virtual audio component refers to the audio processing and output module integrated into the device, responsible for processing and playing audio information, and is usually connected to other devices or systems. User touch commands may also include the touch location and related device identification information.
[0065] When a user touches the current LED bulb, a touch command is generated. The device responds to the touch command, retrieves the virtual audio component identifier corresponding to the current LED bulb, and then obtains the audio data output by the current virtual audio component.
[0066] For example, suppose a user uses a bulb in a smart home system in a specific scenario. Here is a detailed explanation of the process.
[0067] The user walks into the living room and finds the light emitted by the bulb soft and comfortable. Wanting to use the bulb to play some music, he gently touches the bottom of the light. This touch is detected by the device, and the touch command module immediately generates a user touch command.
[0068] In response to a user's touch command, the device first identifies the virtual audio component identifier of the current bulb. This device identifier can point to an associated audio resource library; for example, the current bulb may be connected to multiple audio effects.
[0069] The device queries the corresponding audio database using the virtual audio component identifier to find the audio data output by the currently active virtual audio component. In this example, the audio returned by the database is "the sound of a drum kit." This audio data can be considered an audio clip with a digital file format (such as WAV or MP3) for playback.
[0070] The device retrieves the "drum sound" from the virtual audio component and calls the playback module to play it. At this point, the user can hear the clear drum sound.
[0071] In one implementation, such as Figure 2 As shown, before S101, there are also S107 to S1010, and S107 to S1010 are as follows:
[0072] S107: Display the bulb layout map corresponding to the virtual audio device; wherein, the virtual audio device includes several virtual audio components.
[0073] In this embodiment, the virtual audio device is an audio-related device in a smart home, typically an integrated device that includes audio playback, processing, and control functions. The virtual audio device comprises several virtual audio components. For example, the virtual audio device could be a piano, a drum kit, or a DJ turntable.
[0074] The device generates a visual layout map, specifically a map showing the arrangement of the bulbs corresponding to the virtual audio devices. This map displays the layout of all virtual audio devices and their corresponding bulbs. Users can view this information on the application interface for easy operation and configuration.
[0075] S108: Obtain the relative position of the bulb.
[0076] The relative position of the light bulb in the room is determined by the internal positioning function of sensors or smart devices. The relative position can be represented by a coordinate system (e.g., X, Y, Z coordinates) or described by distance and direction.
[0077] S109: Determine the virtual audio component identifier corresponding to each of the bulbs based on the relative positions of the bulbs and the correspondence between the preset virtual audio component identifier and the relative position.
[0078] The device determines the virtual audio component identifier corresponding to each bulb light based on the relative position obtained by the user and the pre-defined correspondence between virtual audio component identifiers and their relative positions, combined with layout map analysis. This process may involve searching for matching information in the database and grouping according to device functions.
[0079] S1010: Send the virtual audio component identifier corresponding to each of the bulbs to the corresponding bulb.
[0080] After determining the virtual audio component identifier corresponding to each bulb, the system sends these identifiers to each corresponding bulb via wireless communication (such as Wi-Fi or Bluetooth) to facilitate subsequent audio control and playback.
[0081] In this embodiment, users can flexibly configure and control the audio functions of each bulb, not only allowing them to adjust sound quality and effects but also enhancing the interactive experience of smart home systems. Through a graphical interface, relative position acquisition, and communication with smart devices, users are provided with an intuitive and efficient operating method.
[0082] In one embodiment, the user touch command further includes touch duration information and touch pressure information, and obtains the audio data output by the current virtual audio component, including:
[0083] The current virtual audio component is triggered based on the identifier of the virtual audio component corresponding to the current bulb.
[0084] Based on the touch duration information and the touch pressure information, the audio data output by the current virtual audio component is obtained; wherein, the touch duration information is used to determine the duration corresponding to the audio data, and the touch pressure information is used to determine the loudness corresponding to the audio data.
[0085] In this embodiment, the user touch command also includes touch duration information and touch pressure information.
[0086] The touch duration information refers to the length of time a user holds their finger on a light bulb. This touch duration information can be used to determine the duration of the action the user intends to perform.
[0087] Touch pressure information refers to the amount of pressure a user applies to the device when touching it. Different levels of pressure can affect audio output; for example, heavy pressure may correspond to higher loudness.
[0088] The device identifies and triggers the corresponding virtual audio component based on the current bulb's identifier. This process may involve communication with the audio device to activate the specified audio output.
[0089] Simultaneously, the system reads the user's touch duration and touch pressure information. For example, if the touch duration information is identified as 1 second, the corresponding audio data playback duration will be determined to be 1 second. If the touch pressure information is identified as medium pressure, it may be interpreted as medium level of audio loudness, affecting the volume setting.
[0090] Based on touch duration and pressure information, the system generates specific audio data. For example, for a 1-second touch duration, the system might select to play a specific 1-second audio clip. Based on moderate pressure, the audio loudness might be set to 50% of the volume output.
[0091] In this embodiment, users can flexibly control audio output through simple touch operations, combined with duration and pressure information. This interaction method not only enhances the user experience but also demonstrates the intelligence level of smart home devices in audio control.
[0092] S102: Obtain audio timing data corresponding to several segments of audio data output by the current virtual audio component.
[0093] The device accesses the virtual audio component based on the indicator light on the bulb and obtains the audio data currently output by the virtual audio component. This data may include several audio segments and their corresponding audio timing data, such as the start and end times and order of audio playback.
[0094] S103: Perform audio embedding processing on several segments of audio data output by the current virtual audio component and the audio timing data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix.
[0095] The device inputs the acquired audio data and its corresponding audio timing data into the audio processing module to perform audio embedding processing. This process analyzes and integrates each audio segment, generates an audio fusion embedding matrix, and extracts the common features of all audio segments.
[0096] Specifically, two arrays can be created: one to store audio data segments and the other to store the corresponding time-series data. Audio processing libraries (such as Librosa or PyDub) are used to extract features from each audio segment, obtaining MFCC (Mel-frequency cepstral coefficients), spectral features, etc. The extracted features are combined into a large matrix. After processing, an audio fusion embedding matrix is formed, representing the overall features of the audio.
[0097] In one implementation, such as Figure 3 As shown, S103 may include S1031 to S1033, and the specific details of S1031 to S1033 are as follows:
[0098] S1031: Obtain the embedding vector corresponding to the audio data and the embedding vector corresponding to the audio timing data.
[0099] Use an embedding model (such as Word2Vec or Doc2Vec) to generate embedding vectors for the audio data. Assume the audio data is a URL or binary data of an audio segment.
[0100] Similarly, time-series data (such as start time and end time) can be converted into embedding vectors. Time-series data can also be represented as simple numerical features, such as the relative values of start and end times.
[0101] S1032: Input the embedding vector corresponding to the audio data into a preset audio feature extraction model to obtain the embedding feature vector corresponding to the audio data.
[0102] Load the pre-trained audio feature extraction model. This step may use machine learning frameworks such as TensorFlow or PyTorch to invoke the model.
[0103] Input the previously generated audio embedding vector into the audio feature extraction model. Assume this vector is a vector of size N.
[0104] The model calculates and outputs an embedded feature vector based on the input audio embedding vector. This vector represents the high-dimensional features of the audio, which are usually low-dimensional.
[0105] S1033: Encode and fuse the embedding feature vector corresponding to the audio data and the embedding vector corresponding to the audio temporal data to obtain the audio fusion embedding matrix.
[0106] Given the arrangement of audio temporal data, this embedded feature vector can form a multi-track audio, which can then be represented by a matrix. To avoid temporal errors, the embedded vector corresponding to this temporal data can be encoded before or after the embedded feature vector, or in a specified field.
[0107] Prepare the audio embedding feature vector and temporal embedding vector for fusion processing. Choose an appropriate fusion algorithm, such as weighted averaging, concatenation, or deep learning methods (e.g., Siamese Network). For example, a simple concatenation method can be used to join the two vectors together; weights (e.g., 0.8 and 0.2) can be assigned to the two vectors for weighting; or non-linear processing can be performed using activation functions such as ReLU.
[0108] Finally, the encoded and fused vector is parsed into a matrix.
[0109] In this embodiment, by generating embedded vectors, extracting features, and fusing audio data, the system can better understand and respond to user interactions, thereby generating a personalized audio experience. The entire process demonstrates the flexibility and intelligence of smart home systems in audio processing.
[0110] S104: Input the audio fusion embedding matrix into the pre-trained audio recognition model to obtain the first target audio data.
[0111] In this implementation, a deep learning framework (such as TensorFlow or PyTorch) is used to load a pre-trained audio recognition model (such as a CNN or RNN model). The generated audio fusion embedding matrix is input into the pre-trained audio recognition model. This model recognizes audio features and outputs the first target audio data, representing the recognized audio content features, which may be a specific tone, rhythm, etc.
[0112] In one implementation, this helps ensure the model has good recognition performance, enabling smart devices to accurately understand and respond to user audio commands, providing a richer user experience. The training process of the audio recognition model is as follows:
[0113] Acquire the audio training data and the audio training data.
[0114] The audio training data is segmented to obtain several audio training data slices.
[0115] Extract a first training data slice and a second training data slice from the audio training data slice; wherein the second training data slice corresponds to the first training data slice.
[0116] The first training data slice is subjected to audio embedding processing to obtain the audio fusion embedding matrix.
[0117] Based on the audio fusion embedding matrix corresponding to the first training data slice, the second training data slice, the preset loss function, and the preset optimization algorithm, the audio recognition model is iteratively trained until the recognition result of the audio fusion embedding matrix corresponding to the first training data slice meets the preset training termination condition, thereby obtaining the pre-trained audio recognition model.
[0118] In this embodiment, audio training data is obtained from a database or file system. This data may be stored in various formats (such as WAV, MP3, etc.). Assume we have 1000 sound samples covering different categories. Ensure the audio data is normalized for subsequent processing, for example, by uniformly converting different formats to a 16kHz, single-channel WAV format.
[0119] Set the segment duration according to application requirements (e.g., 3 seconds per audio training data segment). Iterate through the audio training data and segment each audio file according to the set duration. For example, cutting a 3-second segment from a 10-second audio file will result in approximately 4 segments (10 / 3). Store the segmented results in a data structure (such as an array or list).
[0120] Assign a unique identifier to each slice to index each segment for easy access later. Select the first training data slice (e.g., slice 1) and its corresponding second training data slice (e.g., slice 2) from the slice array, ensuring that they correspond in time or function.
[0121] Choose a suitable audio embedding model, such as a convolutional neural network (CNN) or a Transformer-based model. Input the first slice of training data into the embedding model to extract audio features and generate audio embedding vectors. If multiple slices are involved, multiple audio feature vectors can be combined to construct a fused embedding matrix.
[0122] The extracted audio fusion embedding matrix and the corresponding second training data slices are loaded into the training set. An appropriate loss function (such as cross-entropy loss) and optimization algorithm (such as Adam) are selected to evaluate model performance and update model parameters.
[0123] During the training process, the audio fusion embedding matrix and the second training data slice are used as inputs to iteratively update the model parameters until the recognition results output by the model meet the preset training conditions (such as reaching 80% accuracy).
[0124] Assuming the training model requires 100 epochs, the model's performance is judged by observing changes in the loss value and the improvement in accuracy during the training process.
[0125] Set conditions, such as the training accuracy reaching a preset standard or the loss value no longer decreasing, and consider the training to be terminated.
[0126] S105: The audio data output by the current virtual audio component and the first target audio data are fused to obtain the second target audio data.
[0127] The device merges the audio data output by the current virtual audio component with the first target audio data. This process may include techniques such as mixing and overlaying to generate the second target audio data, which is the final audio that combines the user's touch intent and recognized audio features.
[0128] For example, an audio processing library (such as pydub) can be used to mix the first target audio data with the current audio data. The mix volume can be adjusted based on touch pressure information.
[0129] Slight mixing: If the pressure is 0-30, use 30% of the first target audio.
[0130] Medium blend: Pressure 31-70, then use 50% of the first target audio.
[0131] Strong Mix: Pressure 71-100, then use 80% of the first target audio.
[0132] In one embodiment, S105 may include S1051 to S1053, such as Figure 4 As shown, S1051 to S1053 are as follows:
[0133] S1051: Perform spectrum conversion on the audio data output by the current virtual audio component and the first target audio data to obtain current spectrum data and first spectrum data.
[0134] The audio data output by the current virtual audio component and the first target audio data are subjected to spectrum conversion to obtain the current spectrum data and the first spectrum data.
[0135] The device uses audio processing libraries (such as Librosa or SciPy) to perform spectral analysis on the current audio data and the first target audio data. Specifically, methods such as Short Time Fourier Transform (STFT) can be used.
[0136] S1052: Based on the current spectrum data, the first spectrum data, and the preset spectrum matching algorithm, add the fusion marker corresponding to the current spectrum data to the first spectrum data.
[0137] Choose a suitable spectrum matching algorithm, such as Dynamic Time Warping (DTW), Euclidean distance, or cosine similarity. Use the spectrum matching algorithm to calculate the similarity between the current spectrum data and the first spectrum data, and add the fusion tag corresponding to the current spectrum data to the first spectrum data.
[0138] S1053: Obtain the second target audio data based on the audio data output by the current virtual audio component, the first target audio data, and the first spectrum data.
[0139] Collect the audio data output by the current virtual audio component, the first target audio data, and the processed first spectrum data (including markers). Perform audio synthesis based on the current audio data and the first target audio data.
[0140] During the synthesis process, audio synthesis libraries (such as Pydub) can be used to combine the processed audio signals: the generated synthesized audio is saved as the second target audio data for subsequent playback and output.
[0141] This embodiment achieves a more natural and richer audio interaction experience between the user and the smart speaker. This technology not only improves user satisfaction but also demonstrates the advanced application potential of smart devices in audio processing.
[0142] S106: Play the second target audio data.
[0143] The second target audio data is played through an audio output device (such as a speaker or headphones), allowing users to experience changes in the audio.
[0144] In one implementation, the lighting effects of smart lamps are controlled by analyzing the spectrum data generated from audio data. This data-driven interaction allows users to experience a richer and more intelligent multimedia environment, enhancing their quality of life and entertainment experience. S106 may include S1061 to S1063, such as... Figure 5 As shown, S1061 to S1063 are as follows:
[0145] S1061: Perform spectrum conversion on the second target audio data to obtain second spectrum data.
[0146] Using an audio processing library such as Librosa or SciPy, perform a Short Time Fourier Transform (STFT) on the second target audio data to obtain spectral data. Store the calculated second spectral data for later use.
[0147] S1062: Generate a lighting control command based on the second spectrum data; wherein the lighting control command includes at least bulb flashing parameters and bulb dynamic color parameters.
[0148] Analysis is performed based on the frequency components of the second spectral data. Special attention can be paid to the intensity of certain frequency ranges (such as low and high frequencies) within the spectrum. For example, statistical characteristics such as maximum, average, and standard deviation can be used for evaluation.
[0149] Light flicker parameters are generated based on the intensity of a specific frequency in the spectral data. A simple rule is established: if the intensity exceeds a certain threshold, a higher flicker rate is set.
[0150] If the spectral intensity value is high, the light may be set to a warm color; if the spectral intensity is low, a cool color may be used to represent dynamic changes.
[0151] The blinking parameters and dynamic color parameters are integrated into a single light control command object, which includes the blinking frequency and dynamic color parameters.
[0152] S1063: Send the lighting control command to the bulb lamp.
[0153] Use a suitable communication protocol (such as MQTT, HTTP, etc.) to send the generated lighting control commands to the bulb.
[0154] In this embodiment, in response to a user touch command, the system acquires the audio data output by the current virtual audio component; acquires the corresponding audio timing data; performs audio embedding processing on the audio data and the corresponding audio timing data to obtain an audio fusion embedding matrix; obtains the first target audio data through the audio fusion embedding matrix; fuses the audio data output by the current virtual audio component and the first target audio data to obtain the second target audio data; and plays the second target audio data. This method allows users to achieve audio interaction and control in smart homes with simple touch actions. Through sophisticated audio processing technology, it ensures a rich auditory experience for users, improving their overall user experience.
[0155] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0156] Please see Figure 6 , Figure 6 This is a schematic diagram of an interactive bulb light device based on audio recognition provided in the second embodiment of this application. The included units are used to perform... Figures 1-5 The steps in the corresponding embodiments. Please refer to the details. Figures 1-5 The relevant descriptions in the corresponding embodiments are shown below. For ease of explanation, only the parts relevant to this embodiment are shown. See also... Figure 6 The audio recognition-based interactive bulb light device 6 includes:
[0157] The first processing unit 61 is used to respond to a user touch command and obtain the audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the virtual audio component identifier corresponding to the current bulb;
[0158] Acquisition unit 62 is used to acquire audio timing data corresponding to several segments of audio data output by the current virtual audio component;
[0159] The second processing unit 63 is used to perform audio embedding processing on several segments of audio data output by the current virtual audio component and the audio time sequence data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix.
[0160] The third processing unit 64 is used to input the audio fusion embedding matrix into a pre-trained audio recognition model to obtain the first target audio data.
[0161] The fourth processing unit 65 is used to fuse the audio data output by the current virtual audio component and the first target audio data to obtain the second target audio data;
[0162] Playback unit 66 is used to play the second target audio data.
[0163] Furthermore, the second processing unit is specifically used for:
[0164] Obtain the embedding vector corresponding to the audio data and the embedding vector corresponding to the audio temporal data;
[0165] The embedding vector corresponding to the audio data is input into a preset audio feature extraction model to obtain the embedding feature vector corresponding to the audio data.
[0166] The embedded feature vector corresponding to the audio data and the embedded vector corresponding to the audio temporal data are encoded and fused to obtain the audio fusion embedding matrix.
[0167] Furthermore, the interactive bulb light device based on audio recognition also includes:
[0168] The fifth processing unit is used to acquire the audio training data and the audio training data;
[0169] The sixth processing unit is used to segment the audio training data to obtain several audio training data slices.
[0170] The seventh processing unit is configured to extract a first training data slice and a second training data slice from the audio training data slice; wherein the second training data slice corresponds to the first training data slice;
[0171] The eighth processing unit is used to perform audio embedding processing on the first training data slice to obtain the audio fusion embedding matrix;
[0172] The ninth processing unit is used to iteratively train the audio recognition model based on the audio fusion embedding matrix corresponding to the first training data slice, the second training data slice, a preset loss function, and a preset optimization algorithm until the recognition result of the audio fusion embedding matrix corresponding to the first training data slice meets the preset training termination condition, thereby obtaining the pre-trained audio recognition model.
[0173] Furthermore, the fourth processing unit is specifically used for:
[0174] The audio data output by the current virtual audio component and the first target audio data are subjected to spectrum conversion to obtain current spectrum data and first spectrum data;
[0175] Based on the current spectrum data, the first spectrum data, and the preset spectrum matching algorithm, a fusion marker corresponding to the current spectrum data is added to the first spectrum data;
[0176] The second target audio data is obtained based on the audio data output by the current virtual audio component, the first target audio data, and the first spectrum data.
[0177] Furthermore, the interactive bulb light device based on audio recognition also includes:
[0178] The tenth processing unit is used to display a map showing the arrangement of bulbs corresponding to the virtual audio device; wherein, the virtual audio device includes several virtual audio components;
[0179] The eleventh processing unit is used to obtain the relative position of the bulb lamp;
[0180] The twelfth processing unit is used to determine the virtual audio component identifier corresponding to each of the bulbs based on the relative positions of the bulbs and the correspondence between the preset virtual audio component identifier and the relative position.
[0181] The thirteenth processing unit is used to send the virtual audio component identifier corresponding to each of the bulbs to the corresponding bulb.
[0182] Furthermore, the first processing unit is specifically used for:
[0183] The current virtual audio component is triggered based on the virtual audio component identifier corresponding to the current bulb light;
[0184] Based on the touch duration information and the touch pressure information, the audio data output by the current virtual audio component is obtained; wherein, the touch duration information is used to determine the duration corresponding to the audio data, and the touch pressure information is used to determine the loudness corresponding to the audio data.
[0185] Furthermore, the playback unit is specifically used for:
[0186] The second target audio data is subjected to spectrum conversion to obtain second spectrum data;
[0187] Based on the second spectrum data, a lighting control command is generated; wherein, the lighting control command includes at least bulb flashing parameters and bulb dynamic color parameters;
[0188] Send the lighting control command to the bulb.
[0189] Figure 7 This is a schematic diagram of an interactive bulb light device based on audio recognition provided in the third embodiment of the present invention. Figure 7 As shown, a terminal device 7 in this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70, such as a program for an interactive method for a bulb lamp based on audio recognition.
[0190] When processor 70 executes computer program 72, it implements the steps in the above-described embodiments of the audio recognition-based bulb light interaction method, for example... Figure 1 Steps S101 to S106 are shown. Alternatively, when the processor 70 executes the computer program 72, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 7 The functions of modules 71 to 76 are shown.
[0191] For example, the computer program 72 can be divided into one or more modules / units, which are stored in the memory 71 and executed by the processor 70 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 72 in the audio recognition-based bulb interactive device. For example, the computer program 72 can be divided into a first processing unit, an acquisition unit, a second processing unit, a third processing unit, a fourth processing unit, and a playback unit, with the specific functions of each unit as follows:
[0192] The first processing unit is used to respond to user touch commands and obtain audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the virtual audio component identifier corresponding to the current bulb;
[0193] The acquisition unit is used to acquire audio timing data corresponding to several segments of audio data output by the current virtual audio component;
[0194] The second processing unit is used to perform audio embedding processing on several segments of audio data output by the current virtual audio component and the audio time sequence data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix.
[0195] The third processing unit is used to input the audio fusion embedding matrix into a pre-trained audio recognition model to obtain the first target audio data.
[0196] The fourth processing unit is used to fuse the audio data output by the current virtual audio component with the first target audio data to obtain the second target audio data;
[0197] The playback unit is used to play the second target audio data.
[0198] The audio recognition-based interactive bulb device 7 provided in this application may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that... Figure 7 This is merely an example of an interactive bulb light device based on audio recognition and does not constitute a limitation on interactive bulb light devices based on audio recognition. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the interactive bulb light device based on audio recognition may also include input / output devices, network access devices, buses, etc.
[0199] The processor 70 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0200] The memory 71 can be an internal storage unit of the audio recognition-based interactive bulb device, such as a hard drive or memory of the audio recognition-based interactive bulb device. The memory 71 can also be an external storage device of the audio recognition-based interactive bulb device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the audio recognition-based interactive bulb device. Furthermore, the audio recognition-based interactive bulb device can include both internal and external storage units. The memory 71 is used to store the computer program and other programs and data required by the audio recognition-based interactive bulb device. The memory 71 can also be used to temporarily store data that has been output or will be output.
[0201] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0202] This application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.
[0203] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0204] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.
[0205] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0206] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0207] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0208] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0209] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0210] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for interactive use of a light bulb based on audio recognition, characterized in that, include: Responding to a user touch command, the system acquires the audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the identifier of the virtual audio component corresponding to the current bulb; Obtain audio timing data corresponding to several segments of audio data output by the current virtual audio component; Audio embedding processing is performed on several segments of audio data output by the current virtual audio component and the audio time sequence data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix; The audio fusion embedding matrix is input into a pre-trained audio recognition model to obtain the first target audio data. The audio data output by the current virtual audio component and the first target audio data are fused together to obtain the second target audio data; Play the second target audio data.
2. The bulb light interaction method based on audio recognition as described in claim 1, characterized in that, The audio embedding process, which involves processing several segments of audio data output by the current virtual audio component and the corresponding audio temporal data of the current virtual audio component to obtain an audio fusion embedding matrix, includes: Obtain the embedding vector corresponding to the audio data and the embedding vector corresponding to the audio temporal data; The embedding vector corresponding to the audio data is input into a preset audio feature extraction model to obtain the embedding feature vector corresponding to the audio data. The embedded feature vector corresponding to the audio data and the embedded vector corresponding to the audio temporal data are encoded and fused to obtain the audio fusion embedding matrix.
3. The bulb light interaction method based on audio recognition as described in claim 1, characterized in that, Before inputting the audio fusion embedding matrix into the pre-trained audio recognition model to obtain the first target audio data, the following steps are included: Acquire the audio training data and the audio training data; The audio training data is segmented to obtain several audio training data slices; Extract a first training data slice and a second training data slice from the audio training data slice; wherein the second training data slice corresponds to the first training data slice; The first training data slice is subjected to audio embedding processing to obtain the audio fusion embedding matrix; Based on the audio fusion embedding matrix corresponding to the first training data slice, the second training data slice, the preset loss function, and the preset optimization algorithm, the audio recognition model is iteratively trained until the recognition result of the audio fusion embedding matrix corresponding to the first training data slice meets the preset training termination condition, thereby obtaining the pre-trained audio recognition model.
4. The interactive method for a bulb lamp based on audio recognition as described in claim 1, characterized in that, The step of fusing the audio data output by the current virtual audio component with the first target audio data to obtain the second target audio data includes: The audio data output by the current virtual audio component and the first target audio data are subjected to spectrum conversion to obtain current spectrum data and first spectrum data; Based on the current spectrum data, the first spectrum data, and the preset spectrum matching algorithm, a fusion marker corresponding to the current spectrum data is added to the first spectrum data; The second target audio data is obtained based on the audio data output by the current virtual audio component, the first target audio data, and the first spectrum data.
5. The bulb light interaction method based on audio recognition as described in any one of claims 1 to 4, characterized in that, Before responding to the user's touch command, the following are included: Displays a map showing the arrangement of bulbs corresponding to the virtual audio device; wherein, the virtual audio device comprises several virtual audio components; Obtain the relative position of the bulb; Based on the relative positions of the bulbs and the correspondence between the preset virtual audio component identifiers and their relative positions, the virtual audio component identifiers corresponding to each bulb are determined. Send the virtual audio component identifier corresponding to each of the bulbs to the corresponding bulb.
6. The bulb light interaction method based on audio recognition as described in any one of claims 1 to 4, characterized in that, The user touch commands also include touch duration information and touch pressure information, and the acquisition of audio data output by the current virtual audio component includes: The current virtual audio component is triggered based on the virtual audio component identifier corresponding to the current bulb light; Based on the touch duration information and the touch pressure information, the audio data output by the current virtual audio component is obtained; wherein, the touch duration information is used to determine the duration corresponding to the audio data, and the touch pressure information is used to determine the loudness corresponding to the audio data.
7. The bulb light interaction method based on audio recognition as described in any one of claims 1 to 4, characterized in that, Playing the second target audio data includes: The second target audio data is subjected to spectrum conversion to obtain second spectrum data; Based on the second spectrum data, a lighting control command is generated; wherein, the lighting control command includes at least bulb flashing parameters and bulb dynamic color parameters; Send the lighting control command to the bulb.
8. An interactive bulb light device based on audio recognition, characterized in that, include: The first processing unit is used to respond to user touch commands and obtain audio data output by the current virtual audio component; wherein, the user touch command is generated when the user touches the current bulb, and the user touch command includes at least the virtual audio component identifier corresponding to the current bulb; The acquisition unit is used to acquire audio timing data corresponding to several segments of audio data output by the current virtual audio component; The second processing unit is used to perform audio embedding processing on several segments of audio data output by the current virtual audio component and the audio time sequence data corresponding to the audio data of the current virtual audio component to obtain an audio fusion embedding matrix. The third processing unit is used to input the audio fusion embedding matrix into a pre-trained audio recognition model to obtain the first target audio data. The fourth processing unit is used to fuse the audio data output by the current virtual audio component with the first target audio data to obtain the second target audio data; The playback unit is used to play the second target audio data.
9. An interactive bulb light device based on audio recognition, characterized in that, include: Processor, memory, and computer programs stored in said memory and executable on said processor; When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.