Sound light effect implementation method and device, electronic equipment and storage medium

CN122803129APending Publication Date: 2026-09-22SHANGHAI XIAODU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610912033.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0002]现有智能音箱在接入大语言模型后,由于数据解析延迟、交互反馈断裂及多终端灯效逻辑不统一等问题,会导致交互体验下降

Benefits of technology

[0010]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122803129A_ABST
    Figure CN122803129A_ABST
Patent Text Reader

Abstract

The present disclosure provides a sound box light effect implementation method and device, electronic equipment and storage medium, relates to the technical field of human-computer interaction, and particularly to the fields of artificial intelligence, voice technology and large language model technology. The specific implementation scheme is: identifying a target interaction node currently where the sound box is located; determining a hardware type of a light effect bearing hardware on the sound box; based on the hardware type, performing light effect mapping to the light effect bearing hardware, and controlling the light effect bearing hardware to display a light effect corresponding to the target interaction node, thereby realizing light effect logic unification under different hardware types, reducing the multi-end development adaptation cost, improving the fluency of human-computer interaction, and optimizing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of human-computer interaction technology, and in particular to the fields of artificial intelligence, speech technology, and large language model technology. Specifically, it relates to a method, device, electronic device, and storage medium for implementing lighting effects in audio. Background Technology

[0002] When existing smart speakers are connected to large language models, the interactive experience deteriorates due to issues such as data parsing delays, broken interactive feedback, and inconsistent lighting effect logic across multiple terminals. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for implementing lighting effects in audio.

[0004] According to one aspect of this disclosure, a method for implementing lighting effects in sound is provided, comprising:

[0005] Identify the target interaction node where the speaker is currently located; Determine the hardware type of the lighting effect carrier hardware on the speaker; Based on the hardware type, light effect mapping is performed on the light effect carrying hardware, and the light effect carrying hardware is controlled to display the light effect corresponding to the target interactive node.

[0006] According to another aspect of this disclosure, a device for realizing sound lighting effects is provided, comprising: The recognition module is used to identify the target interaction node where the speaker is currently located; The determination module is used to determine the hardware type of the lighting effect carrier hardware on the speaker; The control module is used to map lighting effects to the lighting effect carrying hardware based on the hardware type, and to control the lighting effect carrying hardware to display the lighting effects corresponding to the target interactive node.

[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect embodiment.

[0008] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the method described in the first aspect embodiment.

[0009] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect embodiment.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart of a method for implementing sound lighting effects according to an embodiment of this disclosure; Figure 2 This is a flowchart of a method for implementing sound lighting effects according to an embodiment of this disclosure; Figure 3 This is a schematic diagram of lighting effect changes in an overall interactive process provided by an embodiment of this disclosure; Figure 4 This is a schematic diagram of a rhythm change waveform provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of the lighting effect presentation corresponding to an interactive node provided in an embodiment of this disclosure; Figure 6 This is a flowchart of a method for implementing sound lighting effects according to an embodiment of this disclosure; Figure 7 This is a logic flowchart of a conventional audio lighting effect implementation method provided in this embodiment of the disclosure; Figure 8 This is a logic flowchart of a method for implementing sound lighting effects according to an embodiment of this disclosure; Figure 9 This is a structural block diagram of a sound lighting effect implementation device provided in an embodiment of this disclosure; Figure 10 A schematic block diagram of an electronic device used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0013] Human-computer interaction (HCI) is the process of information exchange and interaction between humans and computers. It is a discipline that studies the interactive relationship between systems and users. Systems can be various kinds of machines, as well as computerized systems and software.

[0014] Artificial intelligence (AI) is a new technological science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. It attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence.

[0015] Speech technology, including automatic speech recognition and speech synthesis, is a key technology in the computer field. It is the core technology system for realizing human-computer voice interaction in the computer field, aiming to enable machines to have the ability to "hear", "understand" and "speak".

[0016] Large Language Model (LLM) refers to a deep learning model trained on a large amount of text data, which enables the model to generate natural language text or understand the meaning of language text, learn and simulate the complex rules of human language, and achieve text generation capabilities close to human level.

[0017] As a key application of artificial intelligence technology in home scenarios, smart speakers typically involve a series of voice interactions, including wake-up, listening (receiving sound), analysis (thinking), and broadcasting (responding). In screenless smart speakers, due to the lack of screen feedback, lighting effects become the only and most important visual feedback channel, bearing the key functions of conveying device status, reducing user cognitive burden, and improving the certainty of interaction.

[0018] Currently, the implementation solutions for smart speaker lighting effects can be summarized into the following mainstream technical paths: 1. Basic lighting effect scheme based on limited state indication: Early and some mid-to-low-end smart speaker products adopted the simplest state indication logic, usually using single-color or dual-color LED beads, and used basic modes such as constant light, off, fast flashing, and slow flashing to reflect a limited number of device states.

[0019] The typical implementation is as follows: Standby mode: The lighting effect is always on or off; Wake-up state: The lighting effect lights up instantly or flashes once quickly; Listening status: The lighting effect is constantly on or breathing slowly; Error / Network Anomaly: Lights flash rapidly.

[0020] For example, a red-green dual-color indicator light can be used: solid green indicates standby, green off indicates the start of a conversation, and flashing red indicates network configuration failure.

[0021] 2. Event-driven RGB lighting effect solution: As the cost of RGB LED beads decreases, mid-to-high-end smart speakers are beginning to introduce programmable RGB lighting effects, which achieve richer visual feedback through a predefined event-lighting effect mapping table.

[0022] The technical architecture of this type of solution typically includes: Voice event detection module: Based on local keyword detection or cloud-based Automatic Speech Recognition (ASR) results, it identifies standardized events such as wake-up, playback, pause, error, text-to-speech (TTS) start, and TTS end. Lighting effect mode database: Pre-stores lighting effect parameters for each event, including color (RGB value), animation form (breathing, blinking, flowing, gradient, etc.), duration, brightness curve, etc. Lighting effect driver control module: retrieves lighting effect parameters by looking up a table based on the event type, thereby controlling the output of the LED beads; Finite state machine: manages the switching logic between lighting effect states and avoids state conflicts.

[0023] 3. Lighting effect scheme based on continuous dialogue process: The lighting effect design is more refined.

[0024] In some solutions, a unified light ring is used as the hardware carrier, and all lighting effect solutions can be 100% reused; the process is clearly divided: wake-up → listen → think → respond → end; the use of colors is restrained: usually only blue and orange tones are used, and red or orange is always on when the microphone is disabled.

[0025] In other solutions, a hardware layout with white LEDs arranged in a linear fashion is used; the listening phase after wake-up is further subdivided into two sub-states: "standby" and "receiving"; the lighting effects convey state information through changes in brightness and rhythm, rather than relying on color switching.

[0026] The existing lighting effect display solutions described above have the following drawbacks: 1. A common problem is the lack of a clear resolution / thinking state: During the waiting period between completing voice recording and starting TTS playback, the voice assistant's indicator light abnormally turns off. The integration of large language models exacerbates this issue: the response time of large models is unpredictable (potentially extending to 3-5 seconds or even longer). In existing technologies, when the resolution time exceeds a preset threshold, the indicator light turns off directly or enters idle mode, making it impossible for users to determine whether the device is still working, leading to a false judgment that the "dialogue has ended," and consequently causing invalid operations such as repeated wake-up calls and repeated questioning.

[0027] 2. The granularity of the lighting effect stages is coarse and incomplete.

[0028] 3. The problem of inconsistent lighting effects across different models is serious. The uniformity of lighting effect design overlaps in different speaker models, causing conflicts in the meaning of the lighting effects.

[0029] 4. Lack of pre-planning for lighting effects in hardware design: In existing technologies, lighting effect design usually starts only after the hardware is finalized. This results in the LED layout not taking into account the need for uniformity in lighting effects, the design being limited by the existing physical layout, making it difficult to achieve optimal lighting effect expression, and the lack of a lighting effect design system that can be reused across models and built from the hardware stage.

[0030] 5. Lack of systematic design specifications and parameter system: Existing technologies lack a clear definition of the entire process state machine for the definition of lighting effects, quantitative specifications for lighting effect parameters at each stage, mapping rules from different hardware forms to unified lighting effect parameters, and design knowledge and code components that can be reused across models.

[0031] 6. The lack of a time-based brightness adaptive mechanism and failure to consider the differences in ambient light at different times of day will cause light pollution when used at night due to excessively bright lighting.

[0032] 7. Lack of graceful exit at the end of the broadcast: In existing technologies, the lighting effects usually turn off immediately after the broadcast ends, lacking a transition and disrupting the overall continuity of the interaction.

[0033] Figure 1 This is a flowchart illustrating a method for implementing sound lighting effects according to an embodiment of this disclosure. Figure 1 As shown, the method includes the following steps: S101, identify the target interaction node where the speaker is currently located.

[0034] In some embodiments, an interaction node refers to a process node in the interaction phase between a user and a speaker, wherein the speaker is a smart speaker device capable of receiving and parsing user commands.

[0035] Optionally, the interaction nodes may include, but are not limited to, five interaction nodes: wake-up, listen, parse, respond, and end.

[0036] Optionally, a state machine can be configured inside the speaker to identify the target interaction node that the speaker is currently in.

[0037] S102, Determine the hardware type of the lighting effect carrier hardware on the speaker.

[0038] In some embodiments, the lighting effect carrier hardware refers to a physical light-emitting device that can present lighting effect feedback according to the user's interaction node instructions, and can be determined according to the attribute information of the audio.

[0039] Optionally, the hardware type may include, but is not limited to: LED rings, LED strips, LED matrix, single LED beads, and LED spots.

[0040] In some embodiments, a light ring can be composed of multiple LEDs arranged in a ring to provide a 360° light effect; a light strip can be composed of multiple LEDs arranged in a straight line to provide a linear light effect; a dot matrix can be composed of multiple LEDs arranged in a matrix to provide a graphic light effect; a single LED is a single LED used to provide light effect cues through color and flashing frequency; and a light spot can be a localized focused light-emitting unit composed of closely arranged LEDs, presented in the form of a blocky light spot.

[0041] S103, based on the hardware type, maps lighting effects to the lighting effect carrying hardware, and controls the lighting effect carrying hardware to display the lighting effects corresponding to the target interactive node.

[0042] Optionally, this embodiment can configure different lighting effects for each interactive node according to different hardware types. For example, under the single LED type, different lighting effects can be configured for different interactive nodes; under the dot matrix type, different lighting effects can be configured for different interactive nodes, and so on, to obtain the correspondence between each interactive node and lighting effects under each hardware type.

[0043] Optionally, after determining the current target interaction node and hardware type, the corresponding lighting effect can be determined by matching the target interaction node and the current speaker hardware type. The lighting effect carrying hardware is then controlled to perform lighting effect mapping according to the lighting effect control signal of the target lighting effect, thereby controlling the lighting effect carrying hardware to display the lighting effect corresponding to the target interaction node, and realizing the lighting effect presentation that matches the current target interaction node.

[0044] In this embodiment, the target interaction node where the speaker is currently located and the hardware type of the hardware carrying the speaker's lighting effect are identified. Based on the target interaction node and hardware type, a matching lighting effect is determined for display, achieving precise synchronization between the lighting effect and the interaction node. This ensures unified lighting effect logic across different hardware types, reduces the cost of multi-terminal development and adaptation, improves the smoothness of human-computer interaction, and optimizes the user experience.

[0045] Based on the above embodiments, Figure 2This is a flowchart illustrating a method for implementing sound lighting effects according to an embodiment of this disclosure. Figure 2 As shown, the method includes the following steps: S201, Identify the target interaction node where the speaker is currently located.

[0046] In this embodiment of the disclosure, the method for implementing step S201 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0047] S202, Determine the hardware type of the lighting effect carrier hardware on the audio system.

[0048] Optionally, at least one of the speaker's model information and identification information can be determined; based on at least one of the model information and identification information, the hardware type of the lighting effect carrying hardware can be determined.

[0049] It is understandable that model information is used to uniquely identify the specifications of audio products, while identification information is used to describe the real-time or inherent attributes of audio equipment. Based on either model information or identification information, the hardware type of the lighting effect can be determined, and a more precise lighting effect can be presented based on the hardware type.

[0050] In this embodiment of the disclosure, the method for implementing step S202 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0051] S203, Determine the abstract interaction semantics of the target interaction node.

[0052] Optionally, the abstract interaction semantics can be the intent information conveyed by the speaker when interacting with the user at different interaction nodes. For example, when the target interaction node is "wake up", the corresponding abstract interaction semantics can be "I am here" to convey to the user that the speaker device has been woken up and can listen to the next instruction at any time.

[0053] In some embodiments, when the target interaction node is "listening", the corresponding abstract interaction semantics can be "I am listening, please speak", which conveys to the user that the current waiting stage of the audio device has changed, while expressing that I am listening to you, without any urging semantics, thus optimizing the user experience.

[0054] In some embodiments, when the target interaction node is "parse", the corresponding abstract interaction semantics can be "I'm trying to think, please don't go away", to convey to the user that the device is currently in the stage of understanding the user's needs.

[0055] In some embodiments, when the target interaction node is "response", the corresponding abstract interaction semantics can be "I speak, you listen; if I lag, please don't leave, wait for me a moment", to convey to the user the intent that the current device is ready to respond and that the user should wait patiently due to possible terminal lag.

[0056] It is understandable that when the target interaction node is "end", ending the interaction between the user and the speaker can be done without abstracting the interaction semantics, or the abstract interaction semantics can be set to "okay" to convey to the user the intention to end the interaction.

[0057] S204, based on hardware type and abstract interaction semantics, performs lighting effect mapping to the lighting effect carrying hardware and determines the lighting effect control parameters of the lighting effect carrying hardware.

[0058] Optionally, the lighting effect requirement information of the target interaction node can be determined based on abstract interaction semantics. The lighting effect requirement information includes at least rhythmic features and dynamic modes.

[0059] In some embodiments, the rhythm feature can be the lighting effect response speed. For example, the rhythm feature of each abstract interaction semantic is set to a fast response, so that the user can perceive the device's fast response in a timely manner. For example, the rhythm feature of each interaction node can be defined as a fast response between 200ms and 500ms.

[0060] In some embodiments, the dynamic mode can be a lighting effect, that is, different abstract interaction semantics correspond to different lighting effects. For example, the abstract interaction semantic corresponding to "wake up" is a crisp and obvious entry, a process of lighting up instantly from off to on; the lighting effect corresponding to "listen" is more gentle than that of "wake up", such as using a low-brightness breathing effect and a dynamic effect that fluctuates with the volume, with the effect becoming brighter as the volume increases; the lighting effect corresponding to "analysis" has a rhythm slightly higher than that of "listen", such as a continuous looping dynamic, expressing the process of trying to analyze and think; the lighting effect corresponding to "respond" is more gentle than that of "listen" and "analysis", such as a slow rhythm and reduced light intensity, reducing visual cognitive load and providing users with an environment to listen and think at the same time; the lighting effect corresponding to "end" can be a linear fade-out, that is, the lighting effect gradually turns off, such as fading out within a preset time of 1 second to show the state of ending.

[0061] In other words, differentiated lighting effects are configured for the abstract semantics of the five interactive nodes: wake-up, listening, parsing, responding, and ending. The wake-up node is illuminated at the moment of most obvious change. The parsing node has a lighting effect with a higher rhythm of change than the listening and responding states to reflect the process of trying to analyze and think. The listening node has a higher rhythm of change than the responding node to ensure that the user can reduce visual perception when the speaker responds, and can clearly determine that the speaker is listening during the listening phase. The ending node gradually turns off the lighting effect to reflect the state of the device leaving the device.

[0062] It is understood that in this embodiment, the lighting effect remains on from the moment the speaker is activated until the device fades out, and different dynamic effects are applied to different interaction nodes, such as... Figure 3 As shown, throughout the entire interaction process from the start of the wake-up to the end of the exit, the lighting effects change continuously and there is no instance of the lights going out.

[0063] In some embodiments, the waveform of the lighting effect rhythm change may include, but is not limited to, constant light, breathing, heartbeat, flashing, and pulse forms, as shown in the example below. Figure 4 As shown, rhythm changes can be mapped to different interactive nodes, allowing users to quickly understand the progress.

[0064] Optionally, the physical attribute information of the lighting effect carrying hardware can also be determined based on the hardware type. It is understood that the hardware type includes types such as light ring, light strip, dot matrix, single LED bead and light spot. The hardware type may contain one or more LED beads arranged in a row. In this embodiment, the physical attribute information can be the number of LED beads corresponding to each hardware type and the arrangement of the LED beads. It can also include attribute information such as the color that the hardware type can present. The accuracy of lighting effect control can be improved based on the physical attribute information.

[0065] Furthermore, the lighting effect requirement information and physical attribute information can be mapped to the lighting effect carrying hardware to determine the lighting effect control parameters of the lighting effect carrying hardware; that is, the lighting effect requirement information and physical attribute information are mapped to the lighting effect carrying hardware to obtain the lighting effect control parameters.

[0066] Specifically, based on lighting effect requirement information and physical attribute information, the target physical LEDs that need to participate in lighting effect generation and their corresponding control signals can be determined from the LEDs of the lighting effect carrying hardware; based on the target physical LEDs and their corresponding control signals, the lighting effect control parameters of the lighting effect carrying hardware can be determined, thereby improving the accuracy of obtaining the lighting effect control parameters.

[0067] In some embodiments, the target physical LED beads that need to participate in the generation of the lighting effect refer to the LED beads that need to be lit when the lighting effect is presented. The control signals may include, but are not limited to, signals such as the brightness, color and lighting time of the LED beads. Based on the target physical LED beads and the control signals, lighting effect control parameters accurate to the LED beads are obtained, thereby improving the accuracy of lighting effect control based on the lighting effect control parameters.

[0068] Specifically, in this embodiment, lighting effect requirement information and physical attribute information can be input into the hardware abstraction layer. The hardware abstraction layer performs logical abstraction on the lighting effect requirement information to obtain candidate lighting effect control parameters, wherein the candidate lighting effect control parameters include at least the control signals of virtual LEDs in the virtual LED array. The lighting effect requirement information is abstracted into standardized candidate lighting effect control parameters, which can be decoupled from the dependence on specific hardware types. Furthermore, based on the physical attribute information, the candidate lighting effect control parameters are subjected to morphological normalization mapping, so that the candidate lighting effect control parameters can be mapped to each physical LED, thereby determining the target physical LED and the corresponding control signal. This allows users to avoid repeatedly learning the lighting effect language when using devices with different hardware types, solving the fragmentation problem of existing systems that use a single set of logic for different device products.

[0069] The process of determining the target physical LED bead by performing morphological normalization mapping on the candidate lighting effect control parameters based on physical attribute information is as follows: Based on physical attribute information, the logical position of the virtual LED bead in the virtual LED bead array is mapped to the position of the physical LED bead in the lighting effect carrying hardware; based on the position of the physical LED bead, the target physical LED bead is determined, and the target physical LED bead is the LED bead that participates in the generation of the lighting effect.

[0070] Furthermore, the process of determining the control signal corresponding to the target physical LED is as follows: determine the number of physical LEDs based on physical attribute information; adjust the number of virtual LED control signals based on the number of virtual LEDs and the number of physical LEDs to obtain the control signal of the target physical LED.

[0071] For example, assuming the number of virtual LEDs is greater than the number of physical LEDs, a downsampling strategy can be used to adjust the control signal of the virtual LEDs to determine that multiple virtual LEDs correspond to one physical LED. When acquiring the control signal, the signal features of multiple consecutive virtual LEDs can be fused. For example, regarding the brightness in the signal, the maximum or average brightness of multiple virtual LEDs can be obtained as the brightness parameter of the corresponding physical LED to achieve a one-to-one accurate mapping between virtual LEDs and physical LEDs.

[0072] For example, assuming the number of virtual LEDs is less than the number of physical LEDs, an interpolation strategy can be used to adjust the control signal of the virtual LEDs. For instance, for the brightness parameter in the control signal, an interpolated brightness parameter can be obtained based on the brightness parameters of two adjacent virtual LEDs using a linear interpolation method. This interpolated brightness parameter is then assigned to the physical LED corresponding to the middle position of the virtual LED, thus achieving a precise one-to-one mapping between the virtual LEDs and the physical LEDs.

[0073] S205 controls the lighting effects corresponding to the target interactive node of the hardware display.

[0074] For example, the lighting effects corresponding to each interactive node in this embodiment can be as follows: Figure 5 As shown, different lighting effects can be configured for different hardware types such as single LED beads, LED strips, LED rings, LED spots, and dot matrices.

[0075] Depend on Figure 5 As can be seen, when the hardware type is a single LED, in the "Wake-up" node, the LED lights up immediately, which is a process from off to on. It can flash once or directly enter "Listen" without flashing. In the "Listen" node, the LED breathes, which is relatively slow. In the "Analyze" node, the LED can breathe quickly. If two colors are supported, the two colors can alternate. In the "Response" node, the LED stays on. In the "End" node, the LED slowly turns off.

[0076] When the hardware type is a light strip, in the "Wake Up" node, the light immediately turns on, preferably the entire strip lights up at once, rather than a gradual process; in the "Listen" node, the light strip breathes, which may be accompanied by color changes; in the "Analyze" node, a bright segment slides along the light strip or cycles left and right, or the light strip breathes at a faster speed, alternating between bright and dark; in the "Response" node, the light strip breathes slowly or stays on, with minimal color changes; in the "End" node, it slowly turns off.

[0077] When the hardware type is a light ring, in the "Wake Up" node, the light ring lights up directly, preferably all at once rather than in a gradual transition; in the "Listen" node, the light ring breathes, which can be accompanied by color changes; in the "Analyze" node, a bright segment slides along the light ring or cycles left and right, or the light ring breathes at a faster speed, alternating between bright and dark; in the "Response" node, the light ring breathes slowly or stays on, with minimal color changes; in the "End" node, it slowly turns off.

[0078] When the hardware type is a light spot, at the "Wake Up" node, the light immediately turns on, lighting up the entire area directly, from off to on (brightness 0→1); at the "Listen" node, the light spot breathes, almost flickering, for example, the brightness changes from 0.3→1→0.3, which may be accompanied by color flickering; at the "Resolve" node, the light spot breathes, the brightness is slightly adjusted and the breathing speed is increased, for example, the brightness changes from 0.5→1→0.5; at the "Response" node, the light spot breathes slowly or stays on, with few color changes; at the "End" node, it slowly turns off.

[0079] When the hardware type is dot matrix, in the "Wake Up" node, the dot matrix LEDs light up immediately, with brightness changing from 0 to 1; in the "Listen" node, the dot matrix displays a breathing pattern, flashing when voice is received, with brightness changing from 0.3 to 1 to 0.3; in the "Analyze" node, the dot matrix LEDs alternate between bright and dark, displaying a loading state, with brightness changing from 0.5 to 1 to 0.5; in the "Response" node, the LEDs breathe slowly or remain lit, and can change between single-color brightness; in the "End" node, the LEDs slowly turn off.

[0080] In some embodiments, the "Listen" node can be further subdivided into two states: "Listening" and "Received". The two states can be distinguished, for example, by changing the light effect to breathe during "Listening" or by making the light effect to breathe faster or flash when "Received", so as to more clearly analyze the interaction process and improve the user experience.

[0081] In this embodiment of the disclosure, the method for implementing step S205 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0082] In this embodiment, after determining the target interaction node where the speaker is currently located and the hardware type of the lighting effect carrying hardware, the abstract interaction semantics of the target interaction node are obtained. The lighting effect requirement information of the target interaction node is determined based on the hardware type and the abstract interaction semantics. Physical attribute information is determined based on the hardware type. Based on the lighting effect requirement information and physical attribute information, lighting effect mapping is performed to the lighting effect carrying hardware, thereby determining the lighting effect control parameters and controlling the lighting effect carrying hardware to display the lighting effect corresponding to the target interaction node. This enables unified logic lighting effect display for speaker devices with different hardware types, and different rhythms and frequencies of lighting effect display at different target interaction nodes. Unlike existing technologies that display the same lighting effect at different nodes, this embodiment can convey the real-time operating status of the device to the user through the continuous visual presence of the lighting effect, thus optimizing the user experience.

[0083] Based on the above embodiments, the process for which the target interaction node is the "listen" node will be described. Figure 6 This is a flowchart illustrating a method for implementing sound lighting effects according to an embodiment of this disclosure. Figure 6 As shown, the method includes the following steps: S601 identifies the target interaction node where the speaker is currently located.

[0084] In this embodiment of the disclosure, the method for implementing step S601 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0085] S602, determine the hardware type of the lighting effect carrier hardware on the audio system.

[0086] In this embodiment of the disclosure, the method for implementing step S602 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0087] S603, Determine the abstract interaction semantics of the target interaction node.

[0088] In this embodiment of the disclosure, the method for implementing step S603 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0089] S604, based on hardware type and abstract interaction semantics, maps lighting effects to the hardware that carries the lighting effects, and determines the lighting effect control parameters of the hardware that carries the lighting effects.

[0090] In this embodiment of the disclosure, the method for implementing step S604 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0091] S605 controls the lighting effects of the target interactive node that displays the lighting effects.

[0092] In this embodiment of the disclosure, the method for implementing step S605 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0093] S606, in response to parsing the speech recognition result of the input speaker, controls the speaker to switch from the lighting effect of the speech recognition node to the lighting effect of the model parsing node.

[0094] It should be noted that the voice recognition node is the "listening" node in this embodiment.

[0095] Understandably, during the interaction process of the voice recognition node, the user's voice input commands will be listened to. After receiving the user's voice input commands, the large model configured in the speaker will recognize the input voice. When parsing the recognition results of the voice input to the speaker, the "parse" node will be entered, and the speaker's lighting effects will be switched from the current voice recognition node to the model parsing node.

[0096] For example, suppose the current target interaction node is the "listen" node, which is used for semantic listening and speech recognition. When the speaker parses the speech recognition result, it enters the "parse" node. At this time, the lighting effect of the speech recognition node switches to the lighting effect of the model parsing node, that is, the lighting effect of the "listen" node switches to the lighting effect of the "parse" node.

[0097] In some embodiments, the lighting effect of the model parsing node can be the lighting effect of the "parsing" interactive node in this embodiment, so as to provide feedback to the user that parsing is in progress and optimize the user experience.

[0098] In some embodiments, the rate of streaming character output of the large model can also be determined; the lighting effect control parameters of the model parsing node can be adjusted according to the rate of output characters; for example, for hardware types such as dot matrix, light strip or light ring, the current parsing progress percentage can be determined according to the rate of output characters of the large model, and the progress-increasing lighting effect can be filled on the dot matrix / light strip / light ring according to the percentage of parsing progress. For example, if the parsing output is 10%, then the lighting effect of 10% progress will be displayed on the lighting effect carrying hardware, so that the user can directly determine the current parsing progress according to the lighting effect, and eliminate the user's delay anxiety through "progress visibility".

[0099] In some embodiments, in response to the lighting effect switching to the lighting effect of the model parsing node, an audio player can also be triggered to play audio, wherein the audio playback parameters change synchronously with the lighting effect control parameters of the model parsing node; and / or, a haptic driver can be triggered to generate a vibration signal, wherein the vibration signal is used to provide feedback that the speaker is currently in the model parsing node, and the vibration signal changes synchronously with the lighting effect control parameters of the model parsing node.

[0100] For example, the audio played by the audio player can be a very low-decibel, rhythmic micro-sound effect, such as a slight white noise pulse or a mechanical rotation sound simulation. The sound effect frequency changes synchronously with the light effect frequency, which can supplement the interactive feedback through the auditory channel when the user's eyes are not focused on the speaker, ensuring the continuity of perception throughout the entire chain. The vibration signal can be directed to hardware with a physical feedback mechanism, such as a vibration motor or a physical knob that automatically resets. The micro-vibration feedback provides feedback on the working status of the device, serving as a supplement or replacement for visual feedback. This is suitable for accurately informing visually impaired people or those in environments with strong light interference.

[0101] Optionally, this embodiment can also sense the distance between the user and the speaker, and adjust the lighting effect control parameters of the lighting effect being displayed by the speaker according to the distance, wherein the distance and the lighting effect parameters are positively correlated; for example, the distance between the user and the speaker can be sensed by an infrared sensor or a millimeter-wave radar sensor, and the lighting effect can be dynamically adjusted according to the distance. When the user is far away, a brighter lighting effect is used to ensure visibility, and when the user is close, it automatically switches to an extremely soft lighting effect, further deepening the purpose of scene adaptation, and minimizing the light pollution problem while solving the user's anxiety about the state of mind.

[0102] In some embodiments, the lighting effects being displayed by the speaker can be projected onto the desktop and / or corresponding wall surface where the speaker is located using a projection device on the speaker, so as to convey the information corresponding to the "resolution" node by expanding the light-emitting area, which is more visually immersive.

[0103] S607 monitors the target instructions.

[0104] The target instruction is either a voice playback instruction or a parsing error instruction.

[0105] S608, in response to the absence of a target command, determines that the speaker is still in the model parsing node and continues to display the lighting effects of the model parsing node.

[0106] S609, in response to the detection of a target command, controls the sound to switch from the lighting effect of the model parsing node to the lighting effect of the interaction node associated with the target command.

[0107] In some embodiments, in response to a detected voice playback command, the lighting effect of the speaker is switched from that of the model parsing node to that of the playback node, and the lighting effect of the playback node is made weaker than that of the model parsing node.

[0108] Understandably, after listening to the voice playback command, the system enters the "Response" node and controls the lighting effects of the speakers to switch from the lighting effects of the model analysis node ("Analysis" node) to the lighting effects of the playback node ("Response" node).

[0109] Specifically, the lighting effect control parameters of the playback node differ from those of the parsing node. These lighting effect control parameters may include at least one of the following: lighting effect cycle frequency, lighting effect brightness, and lighting effect rhythm. The lighting effect control parameters of the model's parsing playback node and the playback node must satisfy at least one of the following constraints: The lighting effect loop frequency of the model parsing node is greater than that of the playback node. The lighting effect brightness of the model parsing node is greater than that of the playback node; The lighting effect rhythm of the model parsing node is greater than that of the playback node.

[0110] In some embodiments, in response to the end of the audio interaction and entry into the exit node, the lighting effect of the broadcast node is controlled to linearly decay until the lighting effect is turned off within a set time; that is, when the audio interaction enters the "end" node, the lighting effect is controlled to linearly decay until the lighting effect is turned off within a set time, for example, by fading out linearly over a time of 1-2 seconds, to avoid the feeling of "abruptly stopping" at the end in the prior art. This embodiment simulates the process of light and shadow disappearing in nature, reducing the visual abruptness.

[0111] In some embodiments, in response to the microphone of the speaker being disabled, the ambient light of the speaker is sensed; it is determined that the current ambient light is normal indoor light, dim light, or nighttime; based on the fact that the brightness of the sensed ambient light is less than a set brightness threshold, the brightness of the microphone's disabled warning light effect is reduced; and / or, the disabled color value of the light effect carrying hardware is constrained.

[0112] For example, assuming the current ambient light is dim or at night, that is, when the brightness of the ambient light is less than a set brightness threshold, the brightness of the microphone's disable warning light effect is reduced, and / or the disable color value of the entire device is unified, wherein the microphone's disable warning light can be red or orange; to avoid the light pollution phenomenon of high brightness remaining on after the microphone is disabled in the prior art, this embodiment can reduce the device's interference with the environment.

[0113] In this embodiment, after the target interaction node is a voice recognition node and the lighting effect is accurately presented, in response to the parsing of the voice recognition result, the lighting effect of the current speaker is switched from the voice recognition node to the model parsing node, so that the lighting effect is also displayed in the model parsing node, avoiding the user's mistaken belief that the device has crashed. In the model parsing node state, multi-dimensional signals such as audio and vibration can be used for auxiliary display, more comprehensively ensuring that the user can have full-link status perception. Furthermore, when entering the playback node and ending the node, the lighting effect is displayed according to the corresponding lighting effect strategy, optimizing the user's full-link interactive experience.

[0114] Figure 7 This is a logic flowchart of a conventional audio lighting effect implementation method provided in this embodiment. The user wakes up the device by saying a wake word and enters the listening state. The lighting effect in the listening state is solid blue. Semantic listening and speech recognition are performed in the listening state. At this time, the device waits for the cloud to return the speech result, and the lighting effect is off. After receiving the playback command, the device enters the TTS voice playback state. At this time, the lighting effect flashes periodically within 1 second and turns off after the playback is completed. If there are multiple rounds of dialogue, the device continues to enter the listening state after the TTS voice playback state until the dialogue ends and the lighting effect turns off.

[0115] Based on the above embodiments, Figure 8This is a logic flowchart of a method for implementing sound lighting effects according to an embodiment of the present disclosure. The user says a wake-up word to wake up the device and enter the wake-up state. At this time, the lighting effect gradually changes from off to blue. After entering the wake-up state, the device enters the listening state, where the lighting effect is a constant blue. During the listening state, when voice is detected, the lighting effect is a periodic blue breathing pattern every 1400ms. After the voice monitoring ends and the voice recognition result is received, the device enters the analysis and thinking state, where the lighting effect is a periodic blue breathing pattern every 200ms. After the thinking ends, the device enters the TTS voice broadcasting state, where the lighting effect is a periodic blue breathing pattern every 1600ms. After the TTS broadcast is completed... Once completed, if no further dialogue is needed, the dialogue ends, and the lights slowly dim. In the listening state, if no user speech is detected or no valid instruction is given, the dialogue ends directly. Alternatively, if no valid instruction is given after parsing and a timeout occurs, the dialogue also ends, and the lights slowly dim. In the listening state, if no user speech is detected but a broadcast instruction is received, the dialogue enters the TTS voice broadcast state. In cases of multiple dialogues, the dialogue continues in the listening state after the TTS voice broadcast state for a new round of voice monitoring until the dialogue ends. Corresponding lighting effects are displayed throughout the entire interaction process, optimizing the user experience.

[0116] Figure 9 This is a structural block diagram of a sound lighting effect implementation device provided in an embodiment of this disclosure. For example... Figure 9 As shown, the lighting effect device 900 for the sound system includes: The identification module 901 is used to identify the target interaction node where the speaker is currently located; The determination module 902 is used to determine the hardware type of the lighting effect carrier hardware on the speaker; The control module 903 is used to map lighting effects to the lighting effect carrying hardware based on the hardware type, and to control the lighting effect carrying hardware to display the lighting effects corresponding to the target interactive node.

[0117] In some implementations, control module 903 is used for: Determine the abstract interaction semantics of the target interaction node; Based on hardware type and abstract interaction semantics, lighting effect mapping is performed on the lighting effect carrying hardware to determine the lighting effect control parameters of the lighting effect carrying hardware.

[0118] In some implementations, control module 903 is used for: Based on abstract interactive semantics, determine the lighting effect requirement information of the target interactive node. The lighting effect requirement information includes at least rhythmic features and dynamic modes. The physical attribute information of the hardware that carries the lighting effect is determined based on the hardware type; Based on the lighting effect requirement information and physical attribute information, the lighting effect is mapped to the lighting effect carrying hardware to determine the lighting effect control parameters of the lighting effect carrying hardware.

[0119] In some implementations, control module 903 is used for: Based on the lighting effect requirement information and physical attribute information, the target physical LEDs and corresponding control signals that need to participate in the generation of the lighting effect are determined from the LEDs of the lighting effect carrying hardware. Based on the target physical LED beads and the corresponding control signals, the lighting effect control parameters of the lighting effect carrying hardware are determined.

[0120] In some implementations, control module 903 is used for: Input the lighting effect requirements and physical attribute information into the hardware abstraction layer; The lighting effect requirement information is logically abstracted through the hardware abstraction layer to obtain candidate lighting effect control parameters, which include at least the control signals of virtual LEDs in the virtual LED array. Based on physical attribute information, morphological normalization mapping is performed on the candidate lighting effect control parameters to determine the target physical LED and the corresponding control signal.

[0121] In some implementations, control module 903 is used for: Based on physical attribute information, the logical positions of virtual LEDs in the virtual LED array are mapped to the physical positions of LEDs in the lighting effect carrying hardware. Determine the target physical LED based on its physical LED position; Determine the number of physical LED beads based on physical attribute information; Based on the number of virtual LEDs and the number of physical LEDs, the control signal for the virtual LEDs is adjusted to obtain the control signal for the target physical LEDs.

[0122] In some implementations, module 902 is defined as being used for: Determine at least one of the audio equipment's model information and identification information; The hardware type of the lighting effect carrier hardware is determined based on at least one of the model information and identification information.

[0123] In some implementations, control module 903 is also used for: In response to the parsing of the speech recognition results of the input speaker, the lighting effect of the speaker is switched from the speech recognition node to the model parsing node. Listen for target commands, which may be voice playback commands or parsing error commands. If no target command is detected, and the speaker is still within the model parsing node, the lighting effects of the model parsing node will continue to be displayed; or, In response to the detected target command, the lighting effect of the sound system is switched from the lighting effect of the model parsing node to the lighting effect of the interaction node associated with the target command.

[0124] In some implementations, the control module 903 is also used to perform at least one of the following operations: In response to a detected voice playback command, the lighting effect of the speaker is switched from that of the model analysis node to that of the playback node, and the lighting effect of the playback node is made weaker than that of the model analysis node. In response to the end of the audio interaction and entry / exit node, the lighting effect of the control broadcast node linearly decays until the lighting effect goes out within a set time. In response to the microphone of the speaker being disabled, the system senses the ambient light around the speaker. Based on the perceived ambient light brightness being less than a set brightness threshold, the brightness of the microphone's disable warning light effect is reduced; and / or, Constraints are imposed on the disabled color values ​​of the hardware that carries the lighting effects.

[0125] In some implementations, control module 903 is also used for: Determine the rate of streaming output characters for the large model; The lighting control parameters of the model parsing node are adjusted according to the output character rate.

[0126] In some implementations, the control module 903 is also used to perform at least one of the following operations: In response to the lighting effect switching to the model parsing node, the audio player is triggered to play audio, wherein the audio playback parameters change synchronously with the lighting effect control parameters of the model parsing node; and / or, The haptic driver is triggered to generate a vibration signal, which is used to provide feedback on the current state of the sound system at the model analysis node. The vibration signal changes synchronously with the lighting effect control parameters of the model analysis node.

[0127] In some implementations, the control module 903 is also used to perform at least one of the following operations: The system senses the distance between the user and the speaker, and adjusts the lighting effect control parameters of the lighting effect being displayed by the speaker based on the distance. The distance is positively correlated with the lighting effect parameters. The lighting effects being displayed by the speaker are projected onto the desktop and / or corresponding wall surface where the speaker is located using the projection device on the speaker.

[0128] In this embodiment, after determining the target interaction node where the speaker is currently located and the hardware type of the lighting effect carrying hardware, the abstract interaction semantics of the target interaction node are obtained. The lighting effect requirement information of the target interaction node is determined based on the hardware type and the abstract interaction semantics. Physical attribute information is determined based on the hardware type. Based on the lighting effect requirement information and physical attribute information, lighting effect mapping is performed to the lighting effect carrying hardware, thereby determining the lighting effect control parameters and controlling the lighting effect carrying hardware to display the lighting effect corresponding to the target interaction node. This enables unified logic lighting effect display for speaker devices with different hardware types, and different rhythms and frequencies of lighting effect display at different target interaction nodes. Unlike existing technologies that display the same lighting effect at different nodes, this embodiment can convey the real-time operating status of the device to the user through the continuous visual presence of the lighting effect, thus optimizing the user experience.

[0129] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0130] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0131] Figure 10 A schematic block diagram of an electronic device for implementing embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0132] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0133] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0134] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the method for implementing sound lighting effects. For example, in some embodiments, the method for implementing sound lighting effects can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the method for implementing sound lighting effects described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform a lighting effect implementation method for sound by any other suitable means (e.g., by means of firmware).

[0135] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0136] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0137] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0138] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0139] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0140] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0141] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0142] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for implementing lighting effects in sound, wherein, The method includes: Identify the target interaction node where the speaker is currently located; Determine the hardware type of the lighting effect carrier hardware on the speaker; Based on the hardware type, light effect mapping is performed on the light effect carrying hardware, and the light effect carrying hardware is controlled to display the light effect corresponding to the target interactive node.

2. The method according to claim 1, wherein, The step of mapping lighting effects to the lighting effect-bearing hardware based on the hardware type includes: Determine the abstract interaction semantics of the target interaction node; Based on the hardware type and the abstract interaction semantics, lighting effect mapping is performed on the lighting effect carrying hardware to determine the lighting effect control parameters of the lighting effect carrying hardware.

3. The method according to claim 2, wherein, The step of mapping lighting effects to the lighting effect-bearing hardware based on the hardware type and the interaction semantics, and determining the lighting effect control parameters of the lighting effect-bearing hardware, includes: Based on the abstract interaction semantics, the lighting effect requirement information of the target interaction node is determined, and the lighting effect requirement information includes at least rhythmic features and dynamic modes. The physical attribute information of the hardware carrying the lighting effect is determined based on the hardware type; The lighting effect requirement information and the physical attribute information are mapped to the lighting effect carrying hardware to determine the lighting effect control parameters of the lighting effect carrying hardware.

4. The method according to claim 3, wherein, The process of mapping the lighting effect requirement information and the physical layout information to the lighting effect-bearing hardware to determine the lighting effect control parameters of the lighting effect-bearing hardware includes: Based on the lighting effect requirement information and the physical attribute information, the target physical LEDs and corresponding control signals that need to participate in the lighting effect generation are determined from the LEDs of the lighting effect carrying hardware. Based on the target physical LED beads and the corresponding control signals, the lighting effect control parameters of the lighting effect carrying hardware are determined.

5. The method according to claim 4, wherein, Based on the lighting effect requirement information and the physical attribute information, the process of determining the target physical LEDs and corresponding control signals to participate in the lighting effect generation from the LEDs of the lighting effect carrying hardware includes: Input the lighting effect requirement information and the physical attribute information into the hardware abstraction layer; The lighting effect requirement information is logically abstracted through the hardware abstraction layer to obtain candidate lighting effect control parameters, wherein the candidate lighting effect control parameters include at least the control signals of virtual LEDs in the virtual LED array. Based on the physical attribute information, morphological normalization mapping is performed on the candidate lighting effect control parameters to determine the target physical LED and the corresponding control signal.

6. The method according to claim 5, wherein, The step of performing morphological normalization mapping on the candidate lighting effect control parameters based on the physical attribute information to determine the target physical LED and the corresponding control signal includes: Based on the physical attribute information, the logical positions of the virtual LEDs in the virtual LED array are mapped to the physical LED positions in the lighting effect carrying hardware; The target physical LED bead is determined based on the physical LED bead position; The number of physical LED beads is determined based on the aforementioned physical attribute information; Based on the number of virtual LEDs and the number of physical LEDs, the control signal for the virtual LEDs is adjusted to obtain the control signal for the target physical LEDs.

7. The method according to any one of claims 1-6, wherein, The determination of the hardware type of the lighting effect carrier hardware on the speaker includes: Determine at least one of the model information and identification information of the audio device; The hardware type of the lighting effect carrying hardware is determined based on at least one of the model information and identification information.

8. The method according to any one of claims 1-6, wherein, The method further includes: In response to the parsing of the speech recognition result of the input speaker, the speaker's lighting effect is switched from the speech recognition node to the model parsing node. Listen for target instructions, wherein the target instructions are voice playback instructions or parsing abnormal instructions; In response to the absence of the target command, if it is determined that the speaker is still in the model parsing node, the lighting effects of the model parsing node will continue to be displayed; or, In response to the detected target instruction, the lighting effect of the sound is switched from the lighting effect of the model parsing node to the lighting effect of the interaction node associated with the target instruction.

9. The method according to claim 8, wherein, The method further includes at least one of the following operations: In response to the detected voice playback command, the speaker is controlled to switch from the lighting effect of the model parsing node to the lighting effect of the playback node, and the lighting effect of the playback node is controlled to be weaker than the lighting effect of the model parsing node. In response to the end of the audio interaction and entry into the exit node, the lighting effect of the broadcast node is controlled to linearly decay until the lighting effect is turned off within a set time. In response to the microphone of the speaker being disabled, the ambient light of the speaker is sensed. Based on the perceived ambient light brightness being less than a set brightness threshold, the brightness of the microphone's disable warning light effect is reduced; and / or, The disabled color values ​​of the lighting effect carrying hardware are constrained.

10. The method according to claim 8, wherein, The method further includes: Determine the rate of streaming output characters for the large model; The lighting control parameters of the model parsing node are adjusted according to the rate of the output characters.

11. The method according to claim 8, wherein, The method further includes at least one of the following operations: In response to the lighting effect switching to the lighting effect of the model parsing node, an audio player is triggered to play audio, wherein the playback parameters of the audio change synchronously with the lighting effect control parameters of the model parsing node; and / or, The haptic driver is triggered to generate a vibration signal, wherein the vibration signal is used to provide feedback on the current position of the speaker at the model analysis node, and the vibration signal changes synchronously with the lighting effect control parameters of the model analysis node.

12. The method according to claim 8, wherein, The method further includes at least one of the following operations: The system senses the distance between the user and the speaker, and adjusts the lighting effect control parameters of the lighting effect being displayed by the speaker based on the distance, wherein the distance is positively correlated with the lighting effect parameters; The lighting effects being displayed by the speaker are projected onto the desktop and / or corresponding wall surface where the speaker is located using the projection device on the speaker.

13. A device for achieving sound lighting effects, comprising: The recognition module is used to identify the target interaction node where the speaker is currently located; The determination module is used to determine the hardware type of the lighting effect carrier hardware on the speaker; The control module is used to map lighting effects to the lighting effect carrying hardware based on the hardware type, and to control the lighting effect carrying hardware to display the lighting effects corresponding to the target interactive node.

14. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.

16. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-12.