Multi-mode music rhythm multi-dimensional perception and effect generation display method and system
By using a multi-mode, multi-dimensional perception and effect generation and display method and system for music rhythm, the problems of single rhythm generation and poor synchronization are solved, achieving high-precision rhythm mapping and multi-dimensional effect synchronization, thus enhancing the user's interactive immersion.
Patent Information
- Application Number
- CN202511347371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-01-09
AI Technical Summary
Existing technologies for music rhythm perception, generation, and display suffer from problems such as monotonous rhythm generation, poor synchronization, and a lack of effective interaction, resulting in a weak sense of user immersion.
This invention provides a multi-mode music rhythm multi-dimensional perception and effect generation and display method and system, including a cloud storage module, an edge control module and a human-computer interaction module, which supports custom rhythm mapping, improves the multi-dimensional effect synchronization capability through a delay compensation mechanism, and adds interactive functions.
It improves the accuracy of rhythm mapping, enhances the synchronization and interactivity of multi-dimensional effects, and improves the user's immersive experience.
Smart Images

Figure CN121306068A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and multimedia interactive technology, and in particular to a multi-mode music rhythm multi-dimensional perception and effect generation and display method and system. Background Technology
[0002] As an important art form, music's perception and experience have expanded with the development of multimedia technology. Traditional music rhythm perception mainly relies on a single auditory dimension, supplemented by basic visual feedback (such as voice-controlled lighting). However, as users' demands for immersive experiences increase, existing technologies are gradually revealing significant shortcomings in terms of rhythm flexibility, multi-dimensional effect synchronization, and interactivity.
[0003] Specifically, traditional lighting controls often use fixed algorithms, which cannot customize rhythm mapping and result in insufficient accuracy. They also lack the ability to synchronize multi-dimensional effects (lighting, vibration, etc.) and are prone to timing misalignment between rhythm points and audio / multi-effect subsystems due to transmission delays. Furthermore, they lack interactive matching capabilities, making it impossible to effectively interact with users through rhythm effects, resulting in a weak sense of user immersion.
[0004] Therefore, there is an urgent need for a technical solution that supports custom rhythm mapping and can effectively synchronize multi-dimensional rhythm effects to improve interactivity. Summary of the Invention
[0005] This application provides a multi-mode, multi-dimensional perception and effect generation and display method and system for music rhythm, which solves the technical problems of single rhythm generation, poor synchronization and lack of effective interaction in current music rhythm perception and generation and display technologies, thereby enhancing the user's immersive music experience.
[0006] This application provides a multi-mode, multi-dimensional music rhythm perception and effect generation and display method. This method is applied to a system comprising a cloud storage module, an edge control module, a human-computer interaction module, and a multi-dimensional effect display module. The method includes:
[0007] The system receives a mode selection instruction from the human-computer interaction module to determine the operating mode as either rhythm generation mode or rhythm display mode; wherein, the rhythm generation mode includes manual generation mode, automatic generation mode and semi-automatic generation mode, and the rhythm display mode includes at least the following sub-modes: active display mode, passive display mode and hybrid display mode.
[0008] When the operating mode is rhythm generation mode, rhythm information is generated based on user input and / or system automatic analysis results; wherein, the rhythm information includes at least rhythm point timestamps, control parameters of the corresponding rhythm effect submodule, and effect mapping rules;
[0009] When the operating mode is rhythm display mode, the pre-generated rhythm information is used as the reference rhythm information, and based on the source information of the currently playing audio, the currently playing audio is matched with the reference rhythm information to obtain a first matching result;
[0010] Test data packets are sent to each rhythm effect submodule, and the delay compensation time is calculated based on the detection feedback information corresponding to the test data packets and the pre-trained delay compensation model, so as to adjust the trigger time of each rhythm point in the rhythm information according to the delay compensation time.
[0011] Control commands are sent to each of the rhythm effect sub-modules according to the adjusted trigger time, so as to drive each of the rhythm effect sub-modules to synchronously display the corresponding rhythm effect according to the control commands.
[0012] Compared with the prior art, the significant advantages of this application are as follows:
[0013] This application, through the aforementioned solution, provides users with multiple rhythm generation and display modes, enabling them to customize rhythm mapping, improve rhythm mapping accuracy, and enhance multi-dimensional effect synchronization capabilities through a delay compensation mechanism. It also adds interactive features and improves the user's immersive experience. This effectively solves the technical problems of current music rhythm perception, generation, and display technologies, such as limited rhythm generation and display capabilities, poor synchronization, and lack of effective interaction, thereby enhancing the user's immersive music experience.
[0014] In addition, this application also includes customizable rhythm display device attachment, combination, assembly physical structure (such as lighting, vibration, etc.), and adds a public network platform, which improves the flexibility, scalability and interactivity of rhythm generation. Attached Figure Description
[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0016] Figure 1 This is a flowchart illustrating a multi-mode, multi-dimensional perception and effect generation and display method for music rhythm in an embodiment of this application.
[0017] Figure 2 This is a schematic diagram showing the correspondence between the sensors used to collect motion posture displacement data and body parts in an embodiment of this application.
[0018] Figure 3 This is a schematic diagram illustrating the definition of a rhythm group in a multi-mode music rhythm multi-dimensional perception and effect generation and display method according to an embodiment of this application.
[0019] Figure 4 This is a schematic diagram of audio features in a multi-mode music rhythm multi-dimensional perception and effect generation and display method in an embodiment of this application;
[0020] Figure 5 This is a flowchart illustrating the rhythm adaptation to the musical style and rhythmic atmosphere in the embodiments of this application;
[0021] Figure 6 This is a flowchart illustrating the data processing of the style detection model in this application embodiment;
[0022] Figure 7 This is a schematic diagram of execution delay compensation in a multi-mode music rhythm multi-dimensional perception and effect generation and display method in an embodiment of this application;
[0023] Figure 8 This is a schematic diagram of the rhythm effect information transmission flow in a multi-mode music rhythm multi-dimensional perception and effect generation and display method in an embodiment of this application;
[0024] Figure 9 This is a scene diagram illustrating a multi-mode, multi-dimensional perception and effect generation and display method for music rhythm in an embodiment of this application;
[0025] Figure 10 This is a multimodal generation rhythm-overall architecture diagram in an embodiment of this application;
[0026] Figure 11 This is a schematic diagram of a multimodal generation rhythm-dataflow design structure in an embodiment of this application;
[0027] Figure 12 This is a schematic diagram of visual feature processing and feature mapping in an embodiment of this application;
[0028] Figure 13 This is a schematic diagram illustrating a multimodal generation of rhythmic text description in an embodiment of this application;
[0029] Figure 14 This is a schematic diagram of a fusion mode selection list in an embodiment of this application;
[0030] Figure 15 This is a schematic diagram of the structure of a multi-mode music rhythm multi-dimensional perception and effect generation and display system in an embodiment of this application. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] As users' demands for immersive experiences increase, existing technologies are gradually revealing significant shortcomings in areas such as rhythm flexibility, multi-dimensional effect synchronization, and interactivity. Specifically, traditional lighting controls often rely on fixed algorithms, making it impossible to customize rhythm mapping and resulting in insufficient accuracy. They also lack the ability to synchronize multi-dimensional effects (lighting, vibration, etc.) and are prone to timing misalignment between rhythm points and audio / multi-effect subsystems due to transmission delays. Furthermore, the lack of interactive matching capabilities prevents effective rhythmic interaction with users, leading to a weaker sense of immersion.
[0033] Based on this, the embodiments of this application provide a multi-mode music rhythm multi-dimensional perception and effect generation and display method and system to solve the technical problems of the current music rhythm perception and generation display technology, such as the single rhythm generation and display method, poor synchronization and lack of effective interaction, and to enhance the user's immersive music experience.
[0034] The various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0035] This application provides a multi-mode, multi-dimensional perception and effect generation and display method for music rhythm. This method is applied to a cloud-edge-device collaborative system comprised of a cloud storage module, an edge control module, a human-computer interaction module, and a multi-dimensional effect display module. The modules can communicate with each other via wired or wireless means. Figure 1 As shown, the method may include steps S101-S105:
[0036] S101, the microcontroller receives a mode selection instruction from the human-computer interaction module to determine whether the operating mode is rhythm generation mode or rhythm display mode.
[0037] The rhythm generation mode includes manual generation mode, automatic generation mode and semi-automatic generation mode, and the rhythm display mode includes at least the following sub-modes: active display mode, passive display mode and hybrid display mode.
[0038] It should be noted that the microcontroller is located in the aforementioned edge-end main control module, serving as the execution entity for the multi-mode music rhythm multi-dimensional perception and effect generation display method described in this application. The execution entity can also be other external devices, such as servers, remote servers, server clusters, or edge-end human-computer interaction devices. In simplified systems, the hardware device corresponding to the edge-end main control module can be replaced by a human-computer interaction device, such as a mobile phone, which can perform both control and interaction. That is, the edge-end main control device and the human-computer interaction device can be the same device, capable of performing rhythm generation (e.g., settings via a GUI interface), rhythm transmission, and main control functions. This application does not specifically limit the execution entity. The edge-end main control module is primarily the control center of the edge device, controlling the rhythm effect display sub-module corresponding to the multi-dimensional effect display module to display effects. In active display mode, it receives rhythm information transmitted from the human-computer interaction module and sends it to the multi-dimensional effect display module; in passive display mode, it uses a music sound wave recognition module to align the rhythm effect with the time nodes of the sound. It is also equipped with a display system to show the currently running content. It can generate and display rhythm information in real time by combining the current audio with user configuration and input, and can also perform certain control and operation on the system.
[0039] The cloud storage module is a cloud server data processing and storage system. It primarily provides storage for audio files and automatic generation of rhythm information. It also includes storage for rhythm effects, user information, music files, edge device effect display information, log information, and transmission configuration information. The cloud storage module can also connect to the data processing module, providing not only simple data storage but also information and data processing functions (such as forwarding, matching, and logging). Additionally, it stores configuration information such as device, transmission, and system configurations.
[0040] The human-computer interaction module can be understood as a common human-computer interaction device, such as a mobile phone, computer, tablet, smartwatch, glasses, AR / VR / XR terminal device, etc. It provides a visual operation interface, collects and edits rhythm effect information input by the user, can preview and play rhythm effects, and provides local storage and cloud transmission storage modes. This module mostly involves software system application parts including: a front-end display system (pre-displaying the rhythm with simulated effects); a cloud rhythm / image information database (cloud database) to store rhythm information, video images, music sound information, user information, log file information, configuration information, etc.; and a local rhythm / image information database, etc. The human-computer interaction module applied for in this application can also pull data from cloud storage or store data locally. The system utilizes GUI login and control system interfaces such as APP, QT, LVGL, and web to display the operation interface. This allows for login, acquisition of rhythm list information, display of music-related status, display of effect display device status, display of user profiles, display of log files, reading / uploading of cloud / local data resources, interaction with edge control devices, and setting mapping information between rhythm files and edge effect display subsystems. The corresponding rhythm generation system can generate music-based rhythm effects, such as light effects, impacts, vibrations, rotations, video images, jets, mechanical structure displacement, and rotation. It uses active, automatic, and semi-active methods to generate rhythm effects. Automatically generated rhythm effects are generated from the cloud or locally by interacting with the server, loading pre-made rhythm generation models, combining user preference settings, or using general music rhythm effects.
[0041] The multi-dimensional effect display module is responsible for showcasing rhythmic effects. Through decoupling, each rhythmic effect sub-module is displayed independently. Each rhythmic effect sub-module can also contain different sub-position display points. Different positions can exist within the same effect display module or device. For example, a lighting display device consists of N LEDs, where N is a natural number greater than or equal to 1. Each LED represents a different musical scale or a user-defined effect display position. When defining a rhythm file, different sub-positions can be set at different time periods and different controllable characteristic parameters can be used to achieve different display effects. Each sub-position can be independent or can be combined, cascaded, or linked together to form a whole. If N=1, the effect is displayed as a single position.
[0042] The module can selectively connect to an effects network, providing charging or power module settings, and can feed back key information about the current module status to the control device. Rhythm effects submodules include, for example, sound playback, lighting, vibration and impact, image display, AR / VR / XR rhythm effects display, jetting, and spatial displacement / rotation.
[0043] The sound playback submodule includes one or more audio playback devices used to play music and sound effects from the edge control system. It can control the edge control system by controlling the general interaction module (a human-computer interaction display and control module, a device with at least display and control functions), thereby changing the sound effects, volume, sound characteristics, and audio tracks. It can be connected and organized into multiple modules through expansion or attachment to create more scalable sound effects.
[0044] The lighting submodule includes one or more lighting and 3D rotation control systems, equipped with a 3D rotatable base. By configuring control files, the spotlight beams can be directed in any direction in space, achieving a directional spotlight effect. The lighting system can be controlled, including but not limited to adjusting LED color, brightness, visibility, and display position, to produce effects such as light diffusion range / distance, duration, blur level, fade-in / fade-out, and flickering. The above scenarios represent conventional lighting effects. Lighting pattern displays can be achieved by selecting relevant filters and projections, and can also be controlled by rhythm information. When the lighting effect display uses laser lights, the displayed pattern and color are adjusted by changing cavity vibration displacement, color gating, and the angle of the reflecting mirror; these control information can also be expressed by defining rhythm information. Furthermore, controlling the pattern or related display effects projected by the beam through the light-transmitting sheet, blocking panel, grating pattern, turntable angle, and vibration intensity coordinates can also be defined by rhythm information. When the lighting submodule uses other similar lighting or projection devices, the corresponding effects can also be controlled by defining relevant rhythm information.
[0045] The vibration and percussion submodule contains one or more vibration and percussion sub-position units. For example, when using a two-handed percussion module, each hand can have more than five vibration / percussion units. Each finger can be assigned a motor to simulate percussion effects or a vibration device. According to the central control system, corresponding control signals are assigned to drive these units, such as the intensity of the percussion / vibration, timing, duration of the segment, vibration / percussion sub-position, and turning the vibration / percussion switch on or off at the current position. The locations of the vibration / percussion include, but are not limited to, the inner and outer sides of the fingers, the inner and outer sides of the palm, the inner and outer sides of the wrist, and the areas surrounding the fingers. This effect makes the overall rhythmic effect more three-dimensional and vivid, increases tactile sensation, and enriches the perception of musical rhythm.
[0046] The image display submodule includes one or more display screens for receiving and displaying image or video signals from the edge control system. This includes, but is not limited to, devices with similar effects such as AR / VR / XR devices, smart glasses, and projection devices.
[0047] The lighting submodule also includes a special-shaped lighting strip submodule, used to display corresponding rhythmic effects. It features an integrated, rotatable mechanical structure that allows for adjustments to LED color, brightness, gradient colors, fade-in / fade-out flashing, specific LED positions, and mechanical structure rotation. The AR / VR / XR rhythmic effect display submodule shows the corresponding rhythmic effects.
[0048] The spraying submodule displays spraying effects such as water vapor, dry ice, and flames. It also allows adjustment of spray pressure, direction, length, and opening size, and has functions such as detecting the remaining amount of sprayed material.
[0049] This application also includes a spatial displacement / rotation submodule for mechanical structures: Through the displacement and rotation of motors, mechanical components, and structural parts within space, rhythmic effects are displayed in conjunction with rhythmic information. Furthermore, different rhythmic effect devices can be cascaded or attached. Attachment can be understood as different rhythmic effect submodules working together based on rhythmic information to complete the effect display. For example, a spatial displacement / rotation submodule can be equipped with a lighting display device, and both use general / specific rhythmic information to display the effect. Cascading refers to different effect display devices (in space or when defining a rhythm) being bound together to form a new effect display group.
[0050] The above mode selection command can be issued by the user through the human-computer interaction module (human-computer interaction display and control module), and can select either rhythm generation mode or rhythm display mode. The rhythm generation mode has three different sub-modes, namely active generation mode, passive generation mode and hybrid generation mode.
[0051] S102, when the microcontroller is in rhythm generation mode, it generates rhythm information based on user input and / or system automatic analysis results.
[0052] The rhythm information includes at least the rhythm point timestamp, the control parameters of the corresponding rhythm effect submodule, and the effect mapping rules.
[0053] In this embodiment of the application, when the operating mode is rhythm generation mode, rhythm information is generated based on user input and / or system automatic analysis results, specifically including:
[0054] When the rhythm generation mode is in manual generation mode, the human-computer interaction module receives user input operation signals and parses them to extract rhythm point timestamps, control parameters, and effect mapping rules, generating a custom rhythm file as rhythm information. This can be used in conjunction with a Digital Audio Workstation (DAW) workflow. The operation signals must come from at least one or more of the following: physical buttons, MIDI devices, accelerometers, airflow sensors, touch displays, visual cameras, inertial measurement units, and pressure sensors. In practice, the types of operation signals are not limited to the examples shown here. The effect mapping rules map the rhythm effects corresponding to the control parameters and rhythm point timestamps to the corresponding rhythm effect display submodule and its sub-position mapping method. For example, through a human-computer interaction module with rhythm generation capabilities, users can simulate heptatonic, pentatonic, or other scales and rhythms during music playback by tapping the keyboard with their ten fingers. They can set sub-points mapped to the effect display device, and perform expansion / attenuation / mixing mappings, specifying which effect display sub-module and which sub-point within that module they map to. Alternatively, they can use a method combining sound wave visualization to customize rhythm points one by one. In automatic rhythm generation mode, the currently playing audio is input into a pre-trained rhythm generation model, and the rhythm information is determined based on the model's output (i.e., the system's automatic analysis results). In semi-automatic rhythm generation mode, the rhythm information is determined based on the operation signal and the model's output.
[0055] In the automatic generation mode, the pre-trained rhythm generation model can employ a neural network model. The user inputs a song file (currently playing audio), and the rhythm generation model automatically generates its corresponding rhythm information. It should be noted that the user can configure certain parameters of the model, such as preferences, threshold intensity, and rhythm density. The semi-automatic generation mode fuses the rhythm information obtained from the manual and automatic generation modes. Alternatively, rhythm information can be automatically generated first, and then adjusted by the user through the manual generation mode to obtain the rhythm information for the semi-automatic generation mode. This application does not impose specific limitations on this approach. In the rhythm display mode, the model / configuration parameters can also be used in real-time automatic / semi-automatic mode for rhythm generation, while loading a rhythm effect display (or simulation) device based on temporal alignment and mapping settings.
[0056] In one embodiment of this application, to increase the fun and flexibility of rhythm generation, the method further includes:
[0057] When the operating mode is rhythm generation mode, motion posture displacement data is determined. This motion posture displacement data is acquired through one or more of the following methods: sensor acquisition, motion posture video analysis and generation, and manual control parameter setting. First rhythm information is determined based on the motion posture displacement data. The first rhythm information is input into a preset heterogeneous rhythm effect mapping model to generate second rhythm information corresponding to each currently connected rhythm effect sub-module. The preset heterogeneous rhythm effect mapping model generates a first rhythm vector corresponding to the first rhythm information and converts the rhythm vector into a second rhythm vector corresponding to the respective rhythm effect sub-module according to the module category. The first rhythm vector and each of the second rhythm vectors have a multi-dimensional correlation; this multi-dimensional correlation includes at least mapping relationship type, dimension adaptation rules, and parameter mapping logic. The first rhythm information and the second rhythm information are then added to the rhythm information.
[0058] In other words, this application can collect motion posture displacement data through sensors in a preset dynamic posture sensor group, including but not limited to accelerometers, magnetometers, gyroscopes, visual perception cameras, and LiDAR. For example, a user can wear these sensors while dancing, and motion posture-related data can be extracted from dynamic posture videos (such as a dance video) to obtain motion posture displacement data. Motion posture displacement data can also be obtained by manually setting control parameters. For example, by setting inertial measurement unit (IMU) nodes on body parts to collect motion posture displacement data, such as... Figure 2 As shown, each part of the body corresponds to one or more IMU nodes. Data from various sensors is fused, and through built-in feature extraction and AI algorithms, combined with music sound information, to form a movement posture displacement rhythm effect, i.e., the first rhythm information. When generating this effect, the user needs to use the relevant sensors, or use a device with relevant sensors to perform actions such as moving body parts to allow the sensors to collect relevant data (handheld / carried / tethered / worn-to-the-body devices, etc.), thereby generating the rhythm effect.
[0059] Furthermore, the generated motion effect data can be previewed through motion visualization / virtualization, displayed on visual display devices (such as screens / AR / VR / XR / projectors / glasses), further optimized, or used by users in interactive mode to match the motion posture, displacement, and rhythm effects at corresponding time points using relevant sensor devices. If the rhythm effect matches the defined rhythm effect or the deviation is within a threshold range (defined by the user or system), the rhythm is considered matched, and an evaluation system is used to provide a matching degree and score. Correspondingly, this rhythm also has corresponding prompts in interactive mode before its arrival, such as indicator lights, sounds, and weak vibrations triggered on sensor devices. Successful effects after a single interaction match can also be set, such as generating specific effects like expansion / attenuation / dropout effects on visual display devices or other effect display sub-devices / human-computer interaction devices. Based on weak effect prompts / effect attenuation / effect enhancement and expansion effects, the controllable characteristic parameters of the effect display system / device can be adjusted, and the display can be achieved by adjusting these controllable characteristic parameters. The results can also be uploaded to the cloud or processed locally, awaiting further data processing.
[0060] This application pre-constructs a preset heterogeneous rhythm effect mapping model. This model is a machine learning model trained on several heterogeneous rhythm information samples. After defining the first rhythm information, it can obtain the second rhythm information for different rhythm effect sub-modules. For example, based on the first rhythm information, second rhythm information corresponding to the lighting sub-module, vibration and percussion sub-module, etc., is mapped. Combining the first and second rhythm information, rhythm information corresponding to the music is constructed. The music can be ambient audio, or music pre-stored in the cloud, locally, on a third-party platform, or retrieved from other music playback software via the Internet, etc., without specific limitations. The preset heterogeneous rhythm effect mapping model can map one rhythm information to another. The first rhythm information is not limited to the rhythm information of the above-mentioned action posture displacement data; the first rhythm information can also be the rhythm information corresponding to the lighting sub-module, vibration and percussion sub-module, etc. When converting between rhythm information, sub-position points within the same type of sub-module can also be encoded.
[0061] Specifically, by using a pre-defined heterogeneous rhythm effect mapping model, the first rhythm information can be encoded into a first rhythm vector. Then, by using the module category of each rhythm effect sub-module, the vector transformation matrix and parameters for mutual mapping between different rhythm effect sub-modules can be obtained. By using the vector transformation matrix and the rhythm vector to perform vector calculation, the first rhythm vector can be mapped to other rhythm effect sub-modules. The vector transformation matrix quantifies the multidimensional relationships between vectors, with mapping relationships such as one-to-many (action corresponding to vibration and lighting effects). The dimension adaptation rule adapts each control dimension in the first rhythm vector to the corresponding control dimensions in the second rhythm vector. For example, the first rhythm vector generated from motion posture displacement data corresponds to dimensions such as "motion displacement amplitude, angle change, and speed" (e.g., a 3D vector). When mapped to the rhythm effect submodule of lighting, the second rhythm vector needs to adapt to the control dimensions of the lighting (color RGB / HSV, brightness, on / off switch, i.e., a 4-5D vector). The parameter mapping logic maps the parameter value range in the first rhythm vector to the corresponding parameter value range in the second rhythm vector. For example, the "motion displacement amplitude (e.g., 0-100mm)" in the first rhythm vector is converted into the "vibration intensity (e.g., 0-100%)" of the second rhythm vector in the vibration submodule through parameter mapping logic, i.e., "linear / piecewise function mapping" (e.g., displacement 50mm → intensity 50%). After obtaining the first rhythm information and each of the second rhythm information, the rhythm information is obtained.
[0062] Furthermore, during rhythm effect mapping, there may be changes in the control dimension of a certain rhythm effect sub-module. For example, if the original rhythm effect sub-module of the lights has 5 lights, and two of them cannot be mapped due to some factors, the rhythm effect mapping will be re-executed to transfer the original 5 displayed effects to 3 lights to maintain the rhythm information display effect. The mapping methods can be as follows: 1. Manual mapping (manually set by the user, for example, the original rhythm information is defined as rhythm A, rhythm B, rhythm C, rhythm D, and rhythm E, which are mapped to points 1, 2, 3, 4, and 5 of the light beads respectively. If points 2 and 3 are missing, rhythm B and rhythm C are manually defined and mapped to the remaining points 1, 4, and 5, or these two rhythm effects are directly discarded). 2. Automatic balanced mapping. In more complex cases, this system considers the balance of rhythm effects and performs balanced automatic mapping. It will prioritize mapping the parts of the remaining rhythm points with higher vacancy or lower rhythm density near the time point of the missing point (or through a custom pre-trained model). To ensure the greatest possible rhythmic display effect when using a limited number of points. 3. Or a combination of the above two. That is, semi-automatic. 4. In addition to reducing the mapping of points with the same rhythmic effect, the mapping can also be expanded, for example, N rhythmic effect display points can be expanded to N+X. 5. The point mapping of the same rhythmic effect display device can also be expanded / reduced according to the rhythm group.
[0063] Furthermore, after generating rhythm information, this application may also add rhythm groups to the rhythm information, specifically including:
[0064] The system receives the number of groups N and the grouping criteria in real time. Grouping criteria include at least time interval, rhythm point density, scale type, rhythm point variety, and position point. The rhythm points in the rhythm information are divided into N rhythm groups according to the grouping criteria, and a display intensity coefficient is determined for each rhythm group. The corresponding display intensity coefficient controls the effect display of the rhythm information. The display intensity coefficient corresponds to at least one or more controllable characteristic parameters of the effect display system / equipment, such as light intensity, vibration amplitude, image transparency, spray system intensity, vibration frequency, mechanical component displacement, rotation angle, and spatial displacement rate. These parameters can be set by the user according to the actual usage scenario and are not specifically limited here.
[0065] Rhythm information can be set to multiple rhythm groups. For example... Figure 3 As shown, the definition of a rhythm group is as follows:
[0066] The rhythms appearing at time points A1, A2, and A3 are designated as group A, or the rhythm information appearing within the time interval A1-A2 is designated as group A; the rhythms appearing at time points B1, B2, and B3 are designated as group B, or the rhythm information appearing within the time interval B1-B2 is designated as group B; similarly, group C is designated. When displaying rhythm information in each mode, users can choose to display one / multiple / all rhythm groups. Users can also unlock different rhythm groups gradually, with unlocking conditions including but not limited to reaching a certain level or having specific permissions for specific rhythm effects.
[0067] Rhythm grouping can be done manually, automatically, or semi-automatically. Manual grouping requires the user to select the number of groups and, according to the rhythm group definition, choose the rhythm information to be included in different groups. The user manually assigns the rhythm information. Automatic grouping includes, but is not limited to, random grouping, dense grouping, sparse grouping, tempo-matching grouping, and AI grouping. Random grouping is done randomly by the system; the user only needs to set basic parameters such as the number of groups and percentage (e.g., setting it to divide into two groups, one with 20% rhythm information and the remaining portion with the other group's rhythm effect information). Dense grouping adds more rhythm information to the target group in areas with denser notes and less rhythm information to sparser areas, thus better preserving the rhythmic integrity of the entire song. (For example, if there are 20 rhythmic information points in the 5-10s time period and 120 rhythmic information points in the 15-20s time period, and the user divides the rhythms into two groups, with group A being the primary group for playback (the target group), then the rhythmic information points allocated in the 5-10s period will be relatively fewer, while those in the 15-20s period will be relatively more.) Sparse grouping works in the opposite way. Speed-matched grouping is automatically performed by the system based on the song's local speed, scale changes, timbre composition, and scale density, primarily using speed as a reference, to pre-allocate rhythmic information. Faster time periods are allocated more rhythmic information to the target group, while slower areas are allocated less, achieving a certain match between density and song speed. AI grouping is automatically grouped based on the user-inputted basic parameters and some basic parameters of the song in the speed-matched grouping. Semi-automatic grouping is a combination of the above two methods; the user first performs automatic grouping by adjusting some parameters, and then manually adjusts some rhythm groups or the reverse order to obtain more accurate grouping information.
[0068] In addition, the intensity of rhythm information display for different groups can be set. For example, if there are three groups, A, B, and C, and the user mainly wants to display group A with group B and C as secondary effects, then the intensity of group A can be set to the original intensity, and group B and C can be set to 20%. Alternatively, group B and C can be discarded, i.e., their rhythm display intensity can be set to 0. The intensity distinction of the lighting effect display subsystem is based on light intensity and brightness; the intensity distinction of the vibration impact effect display subsystem is based on vibration frequency, impact speed, and impact force intensity; the intensity of the jet system is based on jet speed, instantaneous compressed gas pressure, and instantaneous speed of the push rod piston; the intensity of the visual dynamic display subsystem is based on motion effects, motion offset, motion amplitude, or effect diffusion amplitude; and the intensity of the mechanical / structural component rotation displacement effect display subsystem is based on displacement distance, rotation angle, and spatial coordinate movement speed. Furthermore, configuration parameters related to intensity, speed, acceleration, quantity, frequency, size, length, range, distance, duration, angle, and measurement can be found in the "Controllable Characteristic Parameters of Effect Display System / Equipment".
[0069] Controllable characteristic parameters of the effect display system / equipment: Controllable and adjustable characteristic parameters of each or every rhythm effect display device or system, including, but not limited to, sound effects, volume, sound effect characteristics, and audio tracks in the sound playback submodule. Lighting submodules include, but are not limited to, spatial displacement orientation, LED color, brightness intensity, display status, pattern effects, light diffusion range / distance, duration, degree of haze (opacity), cavity vibration displacement, color gate, reflective lens angle, light-transmitting sheet, blocking panel, grating pattern, turntable angle, and vibration intensity coordinates. Vibration and impact submodules include, but are not limited to, impact / vibration intensity, time points, vibration impact subunit position, turning the vibration impact switch on or off at the current position, and the position of the vibration impact. Image display submodules include, but are not limited to, image patterns and effects. Spraying submodules include, but are not limited to, spray type, spray pressure, direction, length, and opening size. Spatial displacement / rotation submodules include, but are not limited to, distance, speed, and acceleration in the translational or rotational angle direction within the component structure space.
[0070] S103, when the microcontroller is in rhythm display mode, it uses the pre-generated rhythm information as the reference rhythm information, and performs feature matching between the currently playing audio and the reference rhythm information based on the source information of the currently playing audio to obtain the first matching result.
[0071] In addition, rhythm information can also be generated in real time, without specific limitations here.
[0072] In this embodiment of the application, the above steps specifically include:
[0073] When the rhythm display mode is active, the currently playing audio from the cloud-edge-device collaborative system is determined, and the currently playing audio is aligned with the corresponding baseline rhythm information according to the timestamp alignment rules to obtain a first matching result; wherein, the first matching result includes at least a rhythm preview file after the rhythm and audio are aligned. When the rhythm display mode is passively generated, the currently playing audio from the external environment is acquired, audio recognition is performed on the currently playing audio, and the corresponding pre-associated rhythm information is loaded from the cloud storage module or local storage module according to the audio recognition result to obtain the corresponding first matching result. Audio segments of the currently playing audio are extracted at preset time intervals to compare the audio segments with the audio recognition result. Based on the comparison result, it is determined whether to reload the pre-associated rhythm information and update the first matching result.
[0074] In passive display mode, in addition to matching the song (which song), it is also necessary to match the time point in real time (when the song is playing). If no playing time point is detected (song does not match), it is necessary to rematch the song (which song), or the user can manually enter relevant song information to narrow down the search database.
[0075] The aforementioned pre-associated rhythm information can be stored in a standard audio feature library. This library can pre-store the correspondence between several pieces of music or their features and several pre-made rhythm files (rhythm information). The music or features can be obtained from music playback software, local storage, system cloud, or third-party music media platforms through authorization and processing. Pre-made rhythm files can be obtained through active generation modes, such as manual, automatic, or semi-automatic methods. It should be explained that the standard audio feature library contains music files and music feature information; during playback, the correspondence is found by comparing the music features.
[0076] In addition, the rhythm information used above can also be generated in real time during music playback. By loading a pre-made rhythm generation model and user parameters and preference configurations (which can be shared with the model parameters and methods in the automatic rhythm generation mode), the subsequent mapping (which can be configured by the user or automatically) and loading of the rhythm display device can use the methods in the original mode. No specific restrictions are imposed here.
[0077] When the rhythm display mode is a mixed display mode, the rhythm information corresponding to the first matching result and the music rhythm display information when the audio is played simultaneously are obtained, and the music rhythm display information is matched with the pre-stored rhythm information. In the event that the second matching result is inconsistent, the rhythm point timestamp corresponding to the music rhythm display information is corrected based on the pre-stored rhythm information to dynamically calibrate the playback sequence of the multi-dimensional effect display module and update the first matching result.
[0078] In other words, for currently playing audio from different sources, this application can perform rhythm and music alignment processing on local system music (system source music, where the same source means that both the sound and rhythm come from the system's active playback; the music can come from local / cloud sources), generating a rhythm preview file; it can also collect ambient playback audio, and then find and play the corresponding pre-made rhythm file and align it to the corresponding time node (rhythm point timestamp) in the system's library by comparing the audio-rhythm relationship through features. Furthermore, it uses a preset time interval to detect in real time whether the currently displayed rhythm information corresponds to the ambient audio. This application can simultaneously identify and switch rhythm information. Moreover, it can use a mixed display mode to match and detect whether the currently playing rhythm is consistent with the pre-made rhythm file and whether the timing is aligned. If a song has multiple rhythm effect display files, the user can also choose to configure which rhythm file to use for that song. That is, the rhythm is obtained through sound matching. Thus, during the playback of rhythm information, the music and rhythm are aligned and optimized in real time. The mixed display mode runs simultaneously with at least one of the active display mode and passive display mode. The audio in the mixed mode is actively switched by the user, and the rhythm information is divided into two parts. In the hybrid display mode, rhythm information and execution include: 1. Rhythm information following the system source, playing audio from the same source. 2. Real-time audio detection, determining which song and time point through local, cloud, or third-party platforms, and then matching the rhythm information obtained from the system's rhythm library. 3. Comparing the rhythm information from 1 and 2 to adjust system parameters, such as delay and rhythm groups, to regulate the display effect.
[0079] When matching and detecting environmental audio with pre-associated rhythm information, three different implementation methods can be used. One method uses a neural network for feature generation and feature matching of the pre-made rhythm file, and a sliding window method for feature matching. Another method uses a neural network for both pre-made rhythm file generation and feature matching. The third method uses a non-neural network method for both pre-made rhythm file generation and feature matching. This application does not impose specific limitations on any of these methods.
[0080] For example, a scheme that uses neural networks only for feature generation and feature comparison:
[0081] Cloud preprocessing stage:
[0082] Audio segmentation: Standard audio is divided into smaller frames, each with a 50% overlap. Feature extraction: Each frame is processed using a YAMNet model to extract a 1024-dimensional feature vector. Feature dimensionality reduction: PCA is used to reduce the 1024-dimensional features to 64-dimensional features, and the dimensionality reduction matrix parameters are stored. Storage preparation: The feature matrix is converted into a binary file, generating a feature index file with timestamps corresponding to the parameter information. Model solidification: The model is quantized to INT8 format, and the model structure is optimized to adapt to the neural network processing module of the edge computing device. Resource preloading: The PCA dimensionality reduction matrix is burned into ROM, and the standard audio feature file is stored in Flash memory. Real-time processing at the edge: Audio acquisition: 5 seconds of audio are recorded using a microphone, and pre-emphasis filtering and noise reduction are applied. Feature extraction: The audio is segmented into multiple frame segments. The NPU executes a neural network to extract multi-dimensional features, and the CPU performs PCA dimensionality reduction to a lower feature dimension. Sliding window matching: A sliding window is started from the beginning position of the standard audio features, and the cosine similarity between the segment audio features and the current window is calculated. DTW fine matching selects the three candidate windows with the highest cosine similarity, performs DTW alignment on each candidate window, and calculates the minimum path distance. Decision output: if the minimum DTW distance is less than a threshold, outputs a successful match and its time position; otherwise, outputs a failed match or no match found. Threshold calibration: test audio is collected in a real environment to test the DTW distance distribution and determine the optimal decision threshold.
[0083] For a solution that uses neural networks for both generation and contrast recognition: Data preparation phase: positive sample collection; recording various audio information of the target command; negative sample generation; collecting audio of non-target commands; adding noise; data augmentation; temporal domain augmentation: variable speed (±20%), time shift; noise augmentation: adding multiple types of background noise; frequency domain augmentation: random frequency band masking. Model training phase: feature extraction; samples are processed using an artificial intelligence neural network to extract features and reduce dimensionality; network design; training configuration; setting the loss function, optimizer, batch size, and training epochs. The model is pruned and quantized, and adapted to edge computing devices. Audio processing: audio acquisition, feature extraction, feature aggregation, and Siamese matching are performed step-by-step. Low, medium, and high confidence decision mechanisms are set up and results are output.
[0084] For patterns that use non-neural network methods for feature extraction and matching: In the preprocessing stage, a unified sampling rate is applied: A and B are resampled to the same sampling rate. Noise reduction is then performed: spectral subtraction or Wiener filtering is used to reduce background noise in noisy audio segments. The standard complete segment remains in its original state as a standard reference. In the frame-segmentation stage, the audio is divided into frames of approximately 20-40ms, with a frame shift of 10ms (50% overlap). Feature extraction and detection stage: 13-dimensional MFCC (expandable to higher dimensions, such as first-order and second-order fractal cross-validation) is extracted from each frame. Key parameter configurations include: 40 Mel filters and 2048 FFT points. Audio fingerprint generation is then performed, calculating the spectrogram for each frame and detecting time-frequency peaks: a local window is slid across the spectrogram, retaining the points with the largest amplitudes as keypoints. A hash fingerprint is generated, and keypoint pairs are quantized into integers and combined to form a unique hash value. Feature comparison stage: A fingerprint database of standard audio segments is constructed, storing all fingerprint hash values and corresponding timestamps, and a hash table is used for fast lookup. During the sliding window matching process, the standard audio segment is slid across the database to match the audio segment to be tested. Long windows are used. For each window, a fingerprint set of audio segments is extracted; next, the matching hash value is queried in the fingerprint database of standard audio segments; finally, the consistency of the time offset of the matching fingerprints is statistically analyzed—the time difference between matching pairs is calculated, and a time difference histogram is constructed to find peaks. Decision and post-processing: A judgment is made by setting an empirical threshold; a match is considered successful when the score exceeds the threshold. The threshold can be adjusted using the receiver operating characteristic curve. Non-maximum suppression is used to process adjacent results: adjacent matching windows with a center-to-center distance of less than 1 second are merged, and the highest-scoring matching segment is retained. The final output includes the start / end time of the matching segment in the standard segment; if no match is found, "Not Found" is returned.
[0085] When the operating mode is active display mode, after obtaining the first matching result, music and rhythm can be played synchronously for rhythm preview. Alignment and fine-tuning are performed by adjusting the time difference, start time, and delay. The simulated preview includes various effects such as lighting (light color, brightness, direction, rotation angle, controllable characteristic parameters of the effect display system / device), spotlight spatial turning angle, dynamic effects of the vibration impact module, image content of the video image module, structural component rotation and displacement, spray effects, etc. After verifying the compliance of the rhythm information (e.g., if the rule of only accepting ABC rhythms is violated and other rules require prompting or filtering), the user can choose to upload the file effect information to the cloud for sharing with other users, or keep it locally and load it into the edge control system for playback in the effect display subsystem. Users activate the edge-end effect display device and connect it to the edge-end main control device via wired and / or wireless means. During connection and subsequent transmission, the main control device and the general interactive module (human-computer interaction display and control module) system will display the current transmission status and signal strength, awaiting user operation to map the rhythm signal to the subsystem to be played, and to set rhythm groups, whether to enter interactive mode, set mapping rules, adjust rhythm effects (to make certain adjustments (fine-tuning) before rhythm playback), test and set delay, etc. Through the decoupling design of the subsystem, users can adjust whether the effect display sub-device plays or not.
[0086] In passive display mode, during ambient sound playback, the edge-end main control system collects ambient sound in real time, performs noise reduction and sound extraction, and uses a built-in vector neural network computing module and a pre-trained artificial intelligence model. The collected sound waves are loaded into this system / cloud or a third-party song monitoring system for song identification. After user confirmation, the identified song's sound wave characteristic signal is downloaded to the edge device. Simultaneously, pre-recorded rhythm effects are loaded from the cloud database for synchronized playback (users can simultaneously adjust and configure features such as those described in the active display mode). During subsequent playback, the collected and filtered sound wave segments are compared with standard audio files at regular intervals. If the currently played sound still matches the initially identified sound, playback continues, aligning with time and rhythm nodes. If the difference between the played sound and the initial standard sound characteristics exceeds a threshold, re-identification is performed, and the user is prompted to correct and rematch. After this process, the rhythm information is reloaded, and the effect display subsystem is used for playback and display.
[0087] In the above embodiments, if the rhythm information is generated in real time during music playback rather than pre-made, there is no need to identify the song or compare it with an audio feature-rhythm library. That is, the real-time generated rhythm does not need to be compared with a rhythm library; it can be directly loaded into the effect display device after mapping, user parameter configuration, and other settings.
[0088] In the above embodiments, the musical rhythm information may lack dynamism and cannot be combined with musical style to display rhythmic effects. Therefore, this application also provides the following embodiments, including:
[0089] The audio features corresponding to the rhythm information are determined. These audio features include at least temporal features, frequency domain features, harmonic features, tempo features, and timbre features. The audio features are input into a pre-trained music style detection model to determine the corresponding music style label. Based on the music style label and a preset rhythm group adjustment list, a music style adjustment parameter group corresponding to the rhythm information is determined. The control parameters of each corresponding rhythm effect sub-module are then adjusted according to the music style adjustment parameter group to ensure that the displayed rhythm effect conforms to the stated music style label.
[0090] Specifically, this application incorporates a feature that allows for automatic system-defined or user-defined rhythm and atmosphere during rhythm generation and modification. For example, if the overall rhythm is more relaxed, the edge rhythm effect display on the sub-device will show a softer, more relaxed overall atmosphere when using system-generated or recommended rhythms. Conversely, if the rhythm is faster and stronger, the corresponding rhythm effect will change more rapidly and exhibit more characteristics. The rhythm style serves as an environmental reference factor for rhythm effect generation and modification, and is considered a reference for setting each specific rhythm information point. The overall rhythm of an audio segment can be broken down into different parts, and the rhythm and atmosphere can be set separately for each part.
[0091] Audio feature diagram as follows Figure 4 As shown, the time-domain features include zero-crossing rate, energy envelope, root mean square energy, BPM detection, beat intensity, and rhythm regularity; the frequency-domain features include spectral centroid, spectral bandwidth, and MFCC coefficients; and the harmonic features include chromaticity features and harmonic complexity. Figure 5 A flowchart for adjusting control parameters based on musical style to adapt the rhythm to the rhythmic atmosphere of the music, such as... Figure 5 As shown, by preprocessing or segmenting the audio, conventional features of the aforementioned audio characteristics are extracted, and a music style detection model is used for neural network analysis to classify the music style and obtain music style labels. Furthermore, this application can also add user preferences, adjust the music style based on user preferences, and adjust the audio rhythm information according to the music style and a preset rhythm group adjustment list to make the rhythm information adapt to the rhythmic atmosphere of the music style.
[0092] in, Figure 6 This is a flowchart of the data processing for the style detection model, which includes Mel spectrum input, convolutional layer processing, attention mechanism processing, Bi-LSTM layer processing, fully connected layer processing, style probability output, and other steps.
[0093] In this embodiment, the rhythm display mode further includes an interactive perception mode and a karaoke mode. When the operating mode is interactive perception mode, each rhythm point corresponding to the baseline rhythm information is determined. At a preset time before the timestamp of each rhythm point, a weak effect prompt instruction is generated to cause the corresponding rhythm effect submodule to display the corresponding weak rhythm effect, thus reminding the user to perform an interactive operation. Based on the accuracy of the user's interactive operation triggered by the weak rhythm effect, the rhythm level corresponding to the baseline rhythm information is corrected, and extended overlay effects are added, such as additional effects for celebration or creating a lively atmosphere when the matching degree is high. Different rhythm levels correspond to rhythm information with different rhythm groups. When the operating mode is interactive perception mode, the interactive rhythm matching degree corresponding to the user's interactive operation is determined, and the controllable rhythm variable is adjusted according to the interactive rhythm matching degree to generate a rhythm effect corresponding to the controllable rhythm variable. The rhythm effect includes at least a rhythm extension effect and a rhythm decay effect. The controllable rhythm variable is used to adjust the rhythm information, and the interactive rhythm matching degree is the degree of matching when the user triggers the corresponding operation at a rhythm point.
[0094] In other words, in interactive perception mode, a weak rhythmic effect will be displayed before the rhythm point timestamp to indicate the approaching rhythm. For example, interactive perception mode can include regular interactive mode and live interactive mode. Regular interactive mode is similar to a music rhythm game. For instance, in the rhythm information, at second A, button 1 controls light 1 to display brightness H1, and at second B, button 2 controls light 2 to display brightness H2. At second A or a short time before, light 1 displays brightness J1, which is slightly weaker than H1, or displays an effect in a different way, such as flashing, i.e., a weak rhythmic effect (when displaying sub-effects in a simulated screen manner, it can be indicated by subtlety, text prompts, flashing, etc.). The purpose is to inform the user that an effect will be displayed near the upcoming time point, allowing the user to take action accordingly. For light display effects, weak rhythmic effect prompts include, but are not limited to, reduced light brightness, faster flashing frequency, decreased rotation angle or displacement of mechanical components, and decreased spatial displacement rate, including but not limited to other effect display devices attached to or forming a whole with it, and subtle changes in brightness. For vibration impact effects, weak effect prompts include, but are not limited to, reductions and decreases in vibration intensity and displacement. This also includes reductions and weakening of configuration parameters related to intensity, brightness, speed, acceleration, quantity, frequency, size, length, range, distance, duration, angle, and measurement values in the "Controllable Characteristic Parameters of Effect Display System / Equipment". For other similar effect display sub-modules, weak effect prompts also include, but are not limited to, corresponding effect reduction, fine-tuning of effects, and effect advancement. Additional extended rhythmic effects include, but are not limited to, enhancements of the above effects, or a certain radial pattern, or artistic combinations. Effect decay occurs in the opposite way. If only the on-screen effect information is considered, the method of prompting and displaying the effect is basically similar to common music games. Hitting certain rhythmic notes consecutively will display "wonderful," "perfect," etc., on the screen display device. Alternatively, the controllable characteristic parameters of the effect display system / device may be set too high, or rhythmic effects such as ripples, waves, or spots may be produced, making the overall screen atmosphere more dynamic and exciting – this is expansion. Decay, on the other hand, is relative. If consecutive hits are missed, the atmosphere becomes more monotonous, and the controllable characteristic parameters of the effect display system / device may be set too low (discarding means decay to 0 or not displaying the variable effect). Different preset accuracy thresholds can also be set to match the trigger accuracy, achieving different reduction or expansion rhythmic effects.In interactive mode, enhanced and expanded display effects of the original rhythm effect information can be shown. For example, if the user is at 100% matching level within the time period A1-A2, a more diffuse rhythm effect based on the rhythm near A2 can be displayed for a short period after time point A2. For example, if the original rhythm is to display a ripple effect at the Kth light position, the effect will gradually move to the K-1 and K+1 positions at A2+t1, and at the K-2 and K+2 positions at A2+t1+t2. At this time, this effect can be displayed at a higher frequency, such as activating more lights at the K positions. The preset rhythm discard mode is used to remove rhythm points based on discard rules. The discard rules include at least one of the following: random discard, grouped discard, dense discard, sparse discard, speed matching discard, and AI discard (each discard method is similar to the corresponding grouping method, setting the controllable effect variables of some groups to 0 (or not displaying them) to discard them. The decay rule is also similar to the corresponding discard rule, with decay setting the rhythm to a smaller proportion of the original rhythm intensity). The preset rhythm expansion mode is used to add rhythm information or rhythm groups based on preset expansion rules, while decay does the opposite.
[0095] The effects of the live interactive mode are similar to those of the conventional interactive mode. In this mode, user input and interaction primarily rely on methods such as bullet comments, sent gifts, voice messages, gestures, and image interactions. A short time before the imminent arrival of effect K at a certain point in time (A), a subtle effect, flashing or rotating effect, or corresponding prompts are displayed on the screen or human-computer interaction display device to guide the user to perform appropriate actions based on the rhythmic information. The main control device obtains user input information and the time point by connecting to the live streaming platform's API / SDK to trigger the corresponding effect subsystem for display.
[0096] This application will also analyze the accuracy of user interaction triggers under weak rhythmic prompts to adjust the rhythm level. The higher the rhythm level, the higher the rhythmic point or rhythmic complexity corresponding to the rhythmic information, and vice versa. Rhythmic complexity can be understood as the richness of the quantified rhythmic information display effect.
[0097] Specifically, in the regular interactive mode, user input during sound playback is similar to the manual rhythm generation method, using analog devices such as touch, buttons, playing instruments, IMU, and various sensors. In this mode, real-time user input and standard timing information are displayed and played separately, and the time difference between the two is calculated and recorded in the system. In the live interactive mode, user input and interaction methods primarily involve using bullet comments, sent gifts via voice, and actions / images. Through settings, users can selectively display usernames or user information related to specific interacting users on the screen or in specific locations on the human-computer interaction display device. Furthermore, users can also operate and display effect information from other effect display subsystems in the live interactive mode. This mode incorporates methods for filtering and interacting with specific users. The main control device obtains user interaction information and timing by accessing the live streaming platform's API to trigger the corresponding effect subsystems for effect display.
[0098] In interactive mode, matching results are displayed on the edge control device or edge interactive control and display device through graphical or digital information. Matching accuracy and range precision can be displayed as actual precise numbers or at different levels. For example, if the system defines displaying effect H at time A, and the user triggers the corresponding effect at time A+n, then the error is n, and the accuracy and range precision levels are represented by f(A, n) and g(A, n). Extended rhythm effects can be set based on these two corresponding values. This also includes, but is not limited to, displaying simulated extended information on the corresponding human-computer interaction device. In live interactive mode, in addition to the accuracy and range precision display effects similar to those in regular interactive mode, matching results can also be fed back to different users through replying to bullet comments, responding to graphic animations, and specially designed interactive feedback effect modules, thereby producing customized and unique display effects and enhancing the uniqueness of the interaction. In addition, regular interactive mode can also have responding graphic animations and specially designed interactive feedback effect modules, but only for the current system users.
[0099] Real-time interactive data from the aforementioned interactive modes can be stored locally at breakpoints or uploaded to the cloud via streaming media. When a user finishes a song, corresponding statistical information is displayed. For live interactive modes, this can be displayed through bullet comments, specific screen or human-computer interaction device effects, or specially designed interactive feedback modules. When a user begins to interact again, data analysis can be used to predict the corresponding matching interactive information. Alternatively, the data can be processed in the background at regular intervals by an administrator or specific data processing methods.
[0100] In another embodiment of this application, the method further includes:
[0101] When the operating mode is interactive perception mode, in response to the rhythm practice request from the human-computer interaction module (which can be understood as the practice mode in interactive perception mode), based on the trigger accuracy of the user's interactive operation according to the weak rhythm effect and the preset accuracy threshold range, the corresponding preset rhythm is determined to be discarded / attenuated or the preset rhythm is expanded in order to update the baseline rhythm information.
[0102] In other words, the interactive perception mode of this application also includes a practice mode. The practice mode is used for users to practice rhythm, which can be seen as a process of users gradually adapting to rhythm information. Difficulty levels can be set. In the simple mode, information can be retained for relatively relaxed, slightly slower rhythmic sections, or for strong rhythmic points in fast-paced sections, while discarding dense rhythmic information effects in faster sections. As the difficulty level gradually increases, more rhythmic effect information or groups can be added. The preset accuracy threshold range can be set by the user or developer based on actual usage scenarios, and is not specifically limited here. When the trigger accuracy is within the preset accuracy threshold range or when an operation interaction is successfully performed at the corresponding time rhythm operation point, it can be determined as a preset rhythm expansion. At this time, effects that are expanded or richer than the original preset rhythm effect can be displayed. Conversely, when the accuracy is not within the preset accuracy threshold range or the user does not complete the corresponding operation interaction at the corresponding time rhythm point, it can be determined as a preset rhythm loss / attenuation. At this time, a certain amount of rhythm can be discarded / attenuated. Examples are the same as the examples of rhythm discarding and expansion in the interactive modes mentioned above.
[0103] In addition, the rhythm display mode can also include a karaoke mode. In this mode, the displayed rhythm intensity is related to the matching degree of the singing voice. The edge-end human-computer interaction device collects the current person's singing voice, performs certain noise reduction filtering on the voice, aligns and compares it with the vocal part in the song. At the same time, if the singing matching degree is high within a certain period, additional screen or sub-effect display device effects or extended rhythm effects are displayed; otherwise, discard / attenuation effects are displayed. This is similar to the enhancement and extended display effects in the regular interactive mode. After a song finishes playing, the current karaoke score and comparison with previous scores will be displayed on the image display device. Players can compete with friends, and information such as score rankings can be displayed in real time on the screen of the interactive device subsystem. Data information can be stored locally or uploaded to a cloud database. The overall process includes, but is not limited to, the following key steps: input preprocessing, i.e., detecting the vocal information of the original song, using vocal separation tools such as Spleeter / Demucs to obtain the original vocals and accompaniment. The processed information is stored in the cloud or processed for vocal separation while playing the song. While the user sings, the system processes the captured audio information in real time, performing noise reduction and filtering. If the captured sound contains background music, it continues with vocal separation followed by temporal alignment, such as using Dynamic Time Warping (DTW) alignment. This involves extracting feature information from both the user's and the original song's vocals, calculating the optimal path for alignment, generating an alignment map, and generating the two aligned vocal segments. The system then compares the features of the two segments, including but not limited to timbre, pitch, rhythm, breath control, and emotion. Methods include (but are not limited to) using MFCC cepstral distance, beat interval distribution similarity, energy envelope correlation coefficient, cosine similarity of voiceprint embedding vectors, or neural network methods. Weighted fusion is then performed, assigning different or equal weights to different scores. Each short time interval is used as an evaluation unit, and the score within that interval is calculated. A time-based composite score is used to determine the user's accuracy, and this is fed back to the user for display or appropriate expansion / reduction / dropout effects. The system can record vocals, selectively matching them with song accompaniment, and save them, uploading them to the cloud or storing them on edge devices.
[0104] S104, the microcontroller sends test data packets to each rhythm effect submodule, and calculates the delay compensation time based on the detection feedback information corresponding to the test data packets and the pre-trained delay compensation model, so as to adjust the trigger time of each rhythm point in the rhythm information according to the delay compensation time.
[0105] In this embodiment of the application, before the detection feedback information corresponding to the test data packet and the pre-trained delay compensation model calculate the delay compensation time, the method further includes:
[0106] Acquire historical detection sample data. This historical detection sample data includes at least historical delay records of historical test data packets sent to each rhythm effect submodule, historical preset delay impact parameters, and pre-marked delay compensation times. Input the historical detection sample data into the delay compensation model to be trained to train the model until the model's output accuracy exceeds a predetermined value, thus obtaining a pre-trained delay compensation model.
[0107] In other words, this application has a compensation mechanism that can compensate for the delay in the transmission of rhythm information between the microcontroller and the rhythm effect submodule. Specifically, it uses several historical detection sample data to train a model so as to calculate the delay compensation time for the microcontroller to transmit rhythm information to the rhythm effect submodule in the case of wireless or wired connection.
[0108] The delay compensation time is calculated based on the probe feedback information corresponding to the test data packet and the pre-trained delay compensation model, specifically including:
[0109] Based on the feedback information and the delay calculation formula corresponding to the preset transmission architecture, the corresponding transmission delay is dynamically calculated. The transmission delay and preset delay impact parameters are input into the pre-trained delay compensation model to determine the delay compensation time; the preset delay impact parameters include at least transmission distance, signal strength, and channel utilization.
[0110] Different preset transmission architectures correspond to different delay calculation formulas, such as for Figure 7 The architecture shown, which includes a supernode, uses the following delay calculation formula: t_delay = [(T4-T1)-(T3-T2)] / 2, where t_delay is the transmission delay, T1 is the transmission time from the master controller to the supernode, T2 is the transmission time from the supernode to the ordinary nodes, T3 is the transmission time from the ordinary nodes to the supernode after T2, and T4 is the reception time of the master controller after T3. For a direct connection architecture between the master controller and ordinary nodes, the delay calculation formula can be t_delay = (T4-T1) / 2. When calculating the delay compensation time using the delay compensation model, parameters affecting the delay compensation time can be input. These preset delay impact parameters include at least the current measured delay, segmented average delay (e.g., 1-minute average delay), transmission distance, number of transmission nodes, topology complexity or obstacle influence factor, delay standard deviation, signal strength, channel utilization, retransmission rate, time factor, number of neighboring nodes, chip temperature, network delay, chip system internal delay, and CPU utilization.
[0111] This application also allows manual setting of system delays. Some noticeable delays can be manually set and compensated for, and verified through simulation or playing a short video with a simple rhythm effect. In addition to testing delays and delay compensation, the S104 also includes other user-defined parameter configuration information, such as rhythm groups and playback time offsets, including but not limited to user parameter configuration, rhythm information verification, external device access status, device compliance / activation code verification, rhythm group strength, device availability detection, selective clearing of residual historical garbage cache, pre-made simple test effect demonstrations, mapping scheme selection, and other related initialization, parameter configuration, and device detection functions.
[0112] When the microcontroller and the rhythm effects submodule are far apart, a super node can be set between them, with the rhythm effects submodule acting as a regular node and the microcontroller as the master controller. (Reference) Figure 7 The signal is transmitted to the super nodes in each area via the main transmission path (wired / wireless), and then the super nodes transmit the signal to the sub-devices on nearby ordinary nodes via sub-transmission paths. The functions of the super nodes include, but are not limited to, protocol conversion, data relay, intelligent monitoring and management, physical isolation protection, and encryption authentication protection. A time compensation is required to ensure that the timing of the rhythm information's transmission from the edge control device to the effect display subsystem coincides with the sound playback subsystem's playback of the sound information at the exact moment the song's effect information is defined. The system automatically detects transmission delay upon successful connection initialization and recalculates the delay at regular intervals thereafter. Each time effect information is sent, the displayed time point is the defined standard time point TF(t_delay(calculated delay time)).
[0113] Users can choose from three delay modes: manual, automatic, and semi-automatic. Manual mode allows users to fully set the compensation time, while automatic mode uses the system to automatically calculate the delay time based on the self-tested transmission distance, sending and receiving test distance data packets, and combining this with a preset delay calculation model. Semi-automatic mode combines both, allowing users to set the delay compensation time in conjunction with the automatic mode.
[0114] In wired connections, historical detection sample data includes, but is not limited to, loaded model parameters such as physical distance, transmission distance, network latency, chip system internal latency, network path, number of network devices passed through, RSSI / distance ratio, topology complexity, obstacle influence factor, and historical distance changes, in order to train a latency compensation model (neural network model) related to signal latency. When the device is actually used, the model latency information is inferred and predicted based on the preset parameters at that time and in the early stage, and the transmission latency compensation is obtained.
[0115] In wired connections, if the distance is short, the main controller can directly connect to the rhythm effect submodule via SPI / I2C / RS485 protocols, with almost no need to consider the delay t_delay as (T4-T1) / 2. The main transmission line can, but is not limited to, use a CAN bus, while the sub-transmission lines can use protocols such as RS485 / Ethernet / I2C / SPI. The main and sub-transmission lines can, but are not limited to, use CAN / RS485 / Ethernet / I2C / SPI methods or protocols.
[0116] In the wireless mode, a similar relay super node method can also be used, wherein the transmission between the master controller and the super node, and between the super node and ordinary nodes, includes but is not limited to using Wi-Fi / Bluetooth mesh networks, Zigbee, Thread, LoRa, etc., and the parameter data for training the latency compensation model includes but is not limited to the aforementioned preset latency impact parameters.
[0117] The latency compensation module, once trained, will dynamically calculate the latency compensation time based on the detection feedback information corresponding to the test data packet and the relevant parameters loaded into the model (as mentioned in the previous paragraph). Based on this latency compensation time, the rhythm effect display will be controlled by the sub-device when the rhythm effect is achieved.
[0118] S105, the microcontroller sends control commands to each rhythm effect submodule according to the adjusted trigger time, so as to drive each rhythm effect submodule to synchronously display the corresponding rhythm effect according to the control commands.
[0119] The microcontroller can send timely control commands based on the adjusted trigger time, including but not limited to the time period of each rhythm display, audio, controllable characteristic parameters of the effect display system / device, and sub-position of the effect display system / device, so that each rhythm effect sub-module connected to the microcontroller via wired or wireless means can display rhythm effects. For example, the adjusted trigger time (2.085s) sends a control command to the LED strip in the rhythm effect sub-module ("LEDs 10-20 are on, RGB=255,0,0, brightness 100%, lasting 0.2 seconds"), and at the same time sends a command to the hand vibration unit in the rhythm effect sub-module ("vibration intensity 80%, lasting 0.5s"); the LED strip and the vibration unit respond synchronously to achieve a "visual + tactile" synergistic effect, and the user can view the operating status of each sub-module in real time through a human-computer interaction device (such as LED strip "triggered", remaining battery 80%).
[0120] "Effect Display System / Equipment Sub-positions": Within the same effect display module or device, different positions, such as a lighting display device composed of N LEDs, each representing a different musical scale or user-defined effect display position, can achieve different display effects by being set to different time periods and controlling different controllable characteristic parameters when defining a rhythm file. Each sub-position can be independent or can be combined into a whole through attachment, cascading, or other methods.
[0121] In addition, after generating rhythm information, the method of this application also includes:
[0122] In response to a sharing request from the human-computer interaction module, rhythm information can be sent to a preset public network platform. Based on the update operations of other users on the preset public network platform, a set of secondary creation rhythm information corresponding to the rhythm information is generated. Based on the click rate of each secondary creation rhythm information in the secondary creation rhythm information set, a set of secondary creation rhythm information is generated and displayed to the user through the human-computer interaction module.
[0123] In other words, this application establishes a public network platform to enhance its social attributes. On this platform, users can upload their created effect information files, preview them, like, comment, share, favorite, publish, search, discover, access trending topics, participate in the community, set permissions, block others, send private messages, conduct transactions, express personal preferences, and view push notifications. Users can upload simple rhythm information, images and videos of recorded rhythm information displayed on physical sub-devices, or a combination of both with other text and images. The social platform adds effect fusion attributes and team collaboration functions. Different effect information uploaded by different users can be merged into a richer effect information file. This is similar to gathering different users with verified permissions into a domain or group, where each user has certain permissions to contribute to the creation or modification of some effect information, but only some users have the authority to approve these edits. Methods for previewing the effects include simulation methods (GUI, Web, APP, animation, AR / VR / XR, Unity, etc.), physical sub-device methods, and methods that combine both. Users can publish their own virtual or physical products for trading, such as physical objects with special patterns (e.g., sub-devices displaying effects related to music rhythm), and virtual items such as music and its rhythm display video images and animation effects.
[0124] This application also allows for the customization of the physical architecture or assembly components of rhythm effect display devices, including architectural combinations and attachments of the same or mixed rhythm display devices. For custom rhythm effect display device physical architecture / assemblies: typically, different nodes of the same rhythm display device can be an independent part or partially combined to form different physical structures. For example, in a lighting display effect, there are N lights in total, divided into m groups. Each group plays the same lighting effect. Assuming the first group has k1 lights, the second group has k2 lights, and the m-th group has km lights, then the formula N = k1 + k2 + k3... + km is satisfied. In the m-th group, there are km lights. These lights are fixed together using circuits, wires, and structural components. These hardware, mechanical spatial structures, light positions, decorative shells, assemblies, etc., can all be designed by the user, while the internal code, key circuit components, chips, etc., are still designed by the system design entity.
[0125] For the architecture / assembly of hybrid display equipment: users can customize the combination of different types of rhythm display effect equipment. They can attach, superimpose, or cascade other similar or multiple devices onto a single device (heterogeneous rhythm effect display equipment combination). Attachment: For example, lighting display equipment is attached to a rotating displacement structure device, moving with it, and its display effect will be affected to some extent by the attached device. For example, embedding lighting effect display equipment on a mechanical moving device allows users to customize the structure, hardware, circuitry, etc., forming a single effect display device node. Multiple other similar, heterogeneous, or heterogeneous devices can then form an edge-end effect display system.
[0126] In this application, the user-designed edge-end rhythm effect display device structural assembly, model, circuit connections, physical architecture, surface coating, etc., can be uploaded to a public network platform for sharing, trending, downloading by other users, searching, and other functions similar to defining rhythm information. Simulated rhythm effects and physical video demonstrations can also be displayed on this public network platform.
[0127] In this application, the software components of the edge control device and the human-computer interaction device can be remotely updated via OTA.
[0128] This application can also incorporate a certain random probability function for rhythm information, so that the same music will produce different playback effects based on the existing rhythm, audio, and effect display devices when played at different times. Specifically, before playing the rhythm effect, the user is prompted whether to select this function. If enabled, before each playback, the system automatically identifies the corresponding rhythm points based on the rhythm, audio, and effect display devices, as well as the existing rhythm effect information selected by the user. Random rhythm effect information is generated through a pre-trained model and a certain random probability (data from audio-rhythm correspondence information in a library), or through methods such as spectrum and scale rhythm point recognition.
[0129] In passive display mode, the system first uses a song recognition API to identify the song's fingerprint and determine which song in the song library it belongs to. Then, it determines the specific time point of the song's sequence within the currently playing external system. The function of identifying the song's time point is performed at the edge, using conventional methods or a neural network algorithm embedded in an AI chip for real-time matching and recognition. Deploying the model at the edge aims to reduce latency, as this system requires the song's sound and rhythm to be aligned as closely as possible. However, in this mode, the sound source and rhythm source are not identical (not from the same source). Therefore, to maximize matching efficiency and reduce latency, the model is deployed at the edge using an AI chip. The recognition process uses a pre-trained neural network model, including but not limited to feature extraction, model inference, recognition matching, and output results. The inclusion of a frame-segmentable parallel processing mechanism further enhances the real-time requirements of this design. For example, audio acquisition adopts a parallel processing mechanism of "50ms / frame, 4 frames per group": T=0ms: acquire frame 1; T=50ms: complete frame 1 acquisition, and simultaneously execute "frame 1 feature extraction (NPU) + frame 2 acquisition (microphone)"; T=100ms: complete frame 2 acquisition, and simultaneously execute "frame 2 feature extraction (NPU) + frame 1 model inference (CPU) + frame 3 acquisition (microphone)"; T=150ms: complete frame 3 acquisition, and simultaneously execute "frame 3 feature extraction (NPU) + frame 2 model inference (CPU) + frame 1 result output (CPU) + frame 4 acquisition (microphone)"; Intra-frame / inter-frame optimization: Intra-frame, multi-scale convolution (3×3, 5×5 convolution kernels) are used to process features of different temporal granularities at the same time. Inter-frame, a time-frequency attention layer is used to fuse cross-frame features, reducing the audio recognition time from 200ms to 80ms and improving the response speed of passive mode.
[0130] In addition, when generating and displaying rhythms, this system provides AI-based refining and polishing functions. The AI model provides rhythm information generation or modification suggestions based on the user-selected scene, the amplitude, waveform, phase, duration, timbre, spectrum, dynamic range, spatial characteristics of the music, and the generated information. Users can choose to retain or accept the relevant information, realizing the function of manual and automatic generation and modification.
[0131] Figure 8 This is a schematic diagram of the rhythm effect information transmission flow in the embodiments of this application. The rhythm effect information (rhythm information) can be sent from the cloud or local end to the human-computer interaction device / edge main control device via buffered streaming or pre-loaded full transmission, and then to the edge rhythm effect display device. Alternatively, it can be sent directly from the cloud or local end to the edge rhythm effect display device.
[0132] In this application, latency and latency optimization methods are considered when the edge-end master control device transmits rhythm to the effect display sub-device. This transmission scheme primarily addresses situations where the edge-end effect display sub-device and the human-computer interaction device / edge-end master control device are not on the same global network or obtain rhythms from the same source. That is, the edge-end effect display sub-device needs to relay rhythm effect information from the cloud or locally. However, if the edge-end effect display sub-device can directly obtain rhythm effect information from the cloud via IoT networking protocols such as MQTT or directly connect to a local rhythm / music library, then absolute timestamps (global time) can be used for timing alignment. Regardless of whether the transmission method is local / cloud-based or global / relay, buffered streaming / full preloading methods can be selected. Furthermore, global / relative time fusion alignment is also possible.
[0133] Figure 9 This is a schematic diagram illustrating a scenario of a multi-mode, multi-dimensional perception and effect generation and display method for music rhythm according to an embodiment of this application. It includes a human-computer interaction device / edge-end main control device (microcontroller) sending a request audio to the cloud / local end, and sending an instruction to the edge-end effect display sub-device (rhythm effect sub-module) to cause the rhythm effect sub-module to synchronously request the rhythm. The cloud / local end sends audio data packets and rhythm data packets to the microcontroller and rhythm effect sub-module respectively. The microcontroller performs time alignment, then plays the audio, and the rhythm effect sub-module plays the rhythm. The above operation is repeated cyclically when the buffer is less than a buffer threshold. The buffer threshold can be set by the user or developer according to the actual use case, and is not specifically limited here.
[0134] In addition, this application defines "individual subsystem" as: a system consisting of complete modules such as rhythm control, interaction, sending, sending calculation, receiving, matching, and display.
[0135] The application incorporates individual subsystem networking functionality: Typically, if a user uses an individual subsystem in this application, and wishes to expand it, more and richer rhythm generation or display devices and methods can be used in the rhythm generation or display section. Based on this, other corresponding expansions can be made, such as longer connection lines, higher transmission power, etc. Usually, the various modules of an individual subsystem can be disassembled and separated, and after license verification, reassembled with other devices or modules to form another individual subsystem. In other cases, because an individual subsystem is a fixed set of devices and not easily disassembled, or the original individual subsystem fails to meet the sub-position rhythm mapping relationship in the original mapping, or the user wants to combine / expand / replicate more individual subsystem effects, this application provides an individual subsystem networking function, meaning multiple individual subsystems can be networked to form a more complex network system. The main control of the entire system is undertaken by one of the main control devices or an additional main control device, and the main control of each individual subsystem is undertaken by its internal main control device. When the master controller of an individual subsystem is the main master controller device (primary master controller) of a non-networked system, it receives control signals from the main master controller device and then controls other devices in its individual subsystem. If the master controller device is the master controller device of a certain individual subsystem, it can directly select and allocate rhythms and mappings. Furthermore, the human-computer interaction devices of each individual subsystem can also cooperate to generate rhythms. The rhythm effect display devices of each individual subsystem can be allocated and mapped as a whole, creating more comprehensive rhythm display effects with more or more position points. For example, each vehicle can be considered an individual subsystem, and its front headlights can serve as rhythm display devices. It can then be networked with another vehicle, using a new master controller for systematic control, allocating different headlight positions and rhythm sub-position points to different vehicles, or mapping and expanding different individual subsystems at the same rhythm position point, thereby creating and displaying a more complete, replicated, or expanded rhythm display and effect. The above example illustrates the networking of devices with the same rhythm display effect; different types or mixed display effect devices can also be networked to achieve richer and more systematic effects.
[0136] This application also includes a multimodal (AI)-driven music rhythm perception and generation method. By fusing information from multiple sources such as visual, text, and audio sources, including but not limited to surface mapping (direct correspondence), semantic mapping (conceptual transformation), emotional mapping (sensory transformation), and cross-modal mapping feature fusion, it intelligently generates rhythm markers for the original audio, achieving cross-media artistic expression transformation. The audio A to be generated, along with user-input auxiliary modalities (several images / videos / text / audio B), controllable feature parameters of the effect display system / device, mapping rules, and user settings preferences, are input into the model under this function to generate the rhythm point sequence, mapping information, and configuration parameters of audio A. The process is multimodal feature extraction → cross-modal mapping → rhythm point generation. This method belongs to a type of semi-automatic generation method for rhythm generation.
[0137] See the overall architecture diagram. Figure 10 "Multimodal Rhythm Generation - Overall Architecture Diagram," user input interfaces include, but are not limited to, main audio upload, auxiliary media upload, and parameter configuration. See data flow design. Figure 11 The multimodal rhythm generation dataflow design involves extracting features from the audio to be generated, auxiliary configurations, auxiliary inputs, user configurations, and rhythm mappings (the latter can be extracted separately to obtain a feature vector FX), resulting in a feature vector FA. This FA is then used to form a mapping matrix M, which is fused with FX to form a fused feature F_fusion. This fused feature is then used by the rhythm generation network to generate rhythm information. The main audio analysis stream and... Figure 11 Similarly, visual input processing, text description, and mapping are detailed in the appendix. Figure 12 "Visual Feature Processing and Feature Mapping" Figure 13"Text Description". The feature space mapping process includes, but is not limited to, feature standardization, dimension alignment, attention mechanism, and temporal alignment. Users can choose fusion modes, each with adjustable weights. Modes include, but are not limited to, collaborative mode, where the auxiliary input matches the audio rhythm, and the strategy is to strengthen common feature points, with a default weight of 0.6 for audio and 0.4 for auxiliary input; supplementary mode, where the auxiliary input provides new dimensions, and the strategy is to add rhythm to the blank areas of the audio, with a default weight of 0.7 for audio and 0.3 for auxiliary input; contrast mode, where the condition is to deliberately create a contrast effect, and the strategy is to add visual strong beats to weak beats in the audio, with a default weight of 0.5 for audio and 0.5 for auxiliary input; and overlay mode, where the condition is to rely entirely on the auxiliary input, and the strategy is to replace the original rhythm mode, with a default weight of 0.2 for audio and 0.8 for auxiliary input. The rhythm point generation algorithm includes, but is not limited to, generating a basic grid, hierarchical rhythm generation (primary and secondary beats, decorative rhythms), rhythm point attribute calculation, optimization and smoothing, and humanized processing. Furthermore, adaptive parameter adjustments are possible, such as: audio characteristics → parameter mapping, high rhythmic clarity → reduced auxiliary weights, blurred rhythm → increased auxiliary weights, large dynamic range → increased intensity levels, rich spectrum → refined rhythmic resolution; auxiliary input characteristics → strategy adjustment, obvious visual rhythm → direct mapping, abstract description → emotion-driven, specific instructions → rule execution, etc. The mapping mechanism's layers include, but are not limited to, surface mapping (direct correspondence), semantic mapping (conceptual transformation), and emotional mapping (perception transformation). See attached diagram. Figure 14 "Fusion Mode Selection List"
[0138] Feature understanding and mapping includes, but is not limited to: image understanding and mapping, image understanding layers such as basic layer, semantic layer, sentiment layer, etc., color → music such as hue, brightness, saturation, etc., composition → rhythm such as geometric shape, spatial distribution, line direction, etc.; audio content understanding, audio events such as ambient sound, human voice characteristics, musical style, etc., text semantic understanding such as sentiment mapping, basic sentiment sentiment dimension, natural imagery, action description, spatial description, etc. The above understood and recognized features can be mapped to rhythm (such as tempo, melodic line, loop / flow, regular / staccato, continuous, gradation, etc.), scale (major and minor keys, high and low notes, strong and weak notes, ascending and descending, etc.), rhythm effect mapping rules, etc.
[0139] In addition, to improve the operation of the system and hardware modules, this application also includes the addition of a device and module detection function. For the non-rhythmic effect display portion of the edge control device, general software testing methods can be used, such as unit / integration / system testing, gray-box testing, and general automation / performance testing tools. The rhythmic effect display portion can be detected using a detection device system (consisting of two separate systems with the original system's rhythm generation, display, control, and interaction), including but not limited to sensor detection device modules, parameter display and interaction modules, and detection and analysis control modules. For the rhythmic effect display device, measurement feedback is obtained using standard measurement devices / sensors to obtain measurement data of relevant controllable characteristic parameters in the effect display device, or after combining and calculating data from one or more sensors, this data is compared with the controllable characteristic parameter information originally defined in the system, or visually compared. Combined with user parameter configurations and control device log information, the rhythmic effect display device can be detected and calibrated accordingly. The effect verification and correction of the detection results can be performed locally based on the user manual or relevant libraries, or uploaded to the cloud for further information and technical personnel to operate. This function can be used before the equipment leaves the factory or during user use. The parameter display and interaction module can display measurement parameters, interact with the detection system, and set error thresholds. The detection analysis main control module can measure and fuse information obtained from one or more sensors to analyze and display the controllable characteristic parameters of the system / equipment, and coordinate / feedback to the main control equipment at the edge of the original system for debugging.
[0140] For example, the detection equipment module includes, but is not limited to, the following functions: detecting light, using photodiodes / photoresistors to detect flicker frequency, duration, and rate of change of light intensity; using RGB sensors to detect RGB values, color temperature, and brightness information; and using spectrometers / spectrophotometers to measure spectral distribution. For vibration units, piezoelectric vibration sensors / MEMS accelerometers / laser displacement gauges can be used to detect vibration intensity, and machine vision sensors can be used to measure the shape of light patterns. For mechanical structural components such as motors / motors, Hall effect sensors combined with magnets can measure and calculate displacement, velocity, angle, acceleration, etc.; potentiometers can measure rotation angle; and encoders can measure speed / displacement / rotation speed. For the detection of the injection module, pressure sensors measure injection pressure, flow meters measure and calculate flow rate, ultrasonic / laser methods measure injection distance, and machine vision sensors measure coverage area. It also includes devices that use microphones to detect and calculate sound-related information and machine vision sensors to measure, recognize, and calculate image rhythm display devices.
[0141] This application fully respects the priority and measurement standards of external testing equipment modules. Furthermore, to avoid uncertainties arising from complex system combinations, the testing equipment is primarily used to test the parameters under test, without selecting or designing other complex or superimposed non-essential functions. For example, if a user sets the brightness intensity of a light at time T to A and the color to B in a rhythm file, and uses a rhythm display device to demonstrate the effect, but uses an external testing device where the intensity is A+a and the color is B+b, based on the user-set measurement parameters such as thresholds, if a and b exceed the threshold, there is reason to believe that there is a certain defect in the coordination between the original system (composed of rhythm generation, display, main control, interaction, etc.) and / or rhythm definition / mapping. In this case, the debugging methods and procedures in the aforementioned equipment and module testing functions can be used for corresponding operations and feedback.
[0142] This application, through the aforementioned solution, provides users with multiple rhythm generation and display modes, enabling them to customize rhythm mapping, improve rhythm mapping accuracy, and enhance multi-dimensional effect synchronization capabilities through a delay compensation mechanism. It also adds interactive features and improves the user's immersive experience. This effectively solves the technical problems of current music rhythm perception, generation, and display technologies, such as limited rhythm generation and display capabilities, poor synchronization, and lack of effective interaction, thereby enhancing the user's immersive music experience.
[0143] Figure 15 This application provides a schematic diagram of the structure of a multi-mode music rhythm multi-dimensional perception and effect generation and display system, as shown in the embodiments of this application. Figure 15 As shown, the multi-mode music rhythm multi-dimensional perception and effect generation display system 150 includes:
[0144] The receiving module 151 receives mode selection instructions from the human-computer interaction module to determine whether the operating mode is rhythm generation mode or rhythm display mode. The rhythm generation mode includes manual generation mode, automatic generation mode, and semi-automatic generation mode. The rhythm display mode includes at least the following sub-modes: active display mode, passive display mode, and hybrid display mode. The generation module 152, when the operating mode is rhythm generation mode, generates rhythm information based on user input and / or system automatic analysis results. The rhythm information includes at least rhythm point timestamps, control parameters of the corresponding rhythm effect sub-module, and effect mapping rules. The matching module 153, when the operating mode is rhythm display mode, uses the pre-generated rhythm information as reference rhythm information and performs feature matching between the currently playing audio and the reference rhythm information based on the source information of the currently playing audio to obtain a first matching result. The sending and calculation module 154 sends test data packets to each rhythm effect sub-module and calculates the delay compensation time based on the detection feedback information corresponding to the test data packets and a pre-trained delay compensation model, adjusting the trigger time of each rhythm point in the rhythm information or matching result according to the delay compensation time. The sending module 155 is used to send control commands to each rhythm effect submodule according to the adjusted trigger time, so as to drive each rhythm effect submodule to synchronously display the corresponding rhythm effect according to the control commands.
[0145] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0146] The systems and methods provided in this application are one-to-one correspondences. Therefore, the system also has similar beneficial technical effects as its corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system will not be repeated here.
[0147] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0148] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A multi-mode, multi-dimensional perception and effect generation and display method for music rhythm, applied to a cloud-edge-device collaborative system consisting of a cloud storage module, an edge control module, a human-computer interaction module, and a multi-dimensional effect display module; characterized in that, The method includes: The system receives a mode selection instruction from the human-computer interaction module to determine the operating mode as either rhythm generation mode or rhythm display mode; wherein, the rhythm generation mode includes manual generation mode, automatic generation mode and semi-automatic generation mode, and the rhythm display mode includes at least the following sub-modes: active display mode, passive display mode and hybrid display mode. When the operating mode is rhythm generation mode, rhythm information is generated based on user input and / or system automatic analysis results; wherein, the rhythm information includes at least rhythm point timestamps, control parameters of the corresponding rhythm effect submodule, and effect mapping rules; When the operating mode is rhythm display mode, the pre-generated rhythm information is used as the reference rhythm information, and based on the source information of the currently playing audio, the currently playing audio is matched with the reference rhythm information to obtain a first matching result; Test data packets are sent to each rhythm effect submodule, and the delay compensation time is calculated based on the detection feedback information corresponding to the test data packets and the pre-trained delay compensation model, so as to adjust the trigger time of each rhythm point in the rhythm information according to the delay compensation time. Control commands are sent to each of the rhythm effect sub-modules according to the adjusted trigger time, so as to drive each of the rhythm effect sub-modules to synchronously display the corresponding rhythm effect according to the control commands.
2. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 1, characterized in that, When the operating mode is rhythm generation mode, rhythm information is generated based on user input and / or system automatic analysis results, specifically including: When the rhythm generation mode is the manual generation mode, the human-computer interaction module receives user input operation signals and parses the operation signals to extract the rhythm point timestamps, each of the control parameters, and the effect mapping rules, generating a custom rhythm file as the rhythm information; wherein, the operation signals come from at least one or more of physical buttons, MIDI devices, accelerometers, airflow sensors, touch display devices, vision cameras, and inertial measurement units; the effect mapping rules are a mapping method that maps the rhythm effects corresponding to the control parameters and the rhythm point timestamps to the corresponding effect display rhythm effect submodules; When the rhythm generation mode is the automatic generation mode, the corresponding currently playing audio is input into the pre-trained rhythm generation model to determine the rhythm information based on the model output. When the rhythm generation mode is the semi-automatic generation mode, the rhythm information is determined based on the operation signal and the model output result.
3. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 1, characterized in that, When the operating mode is rhythm display mode, the pre-generated rhythm information is used as the reference rhythm information, and based on the source information of the currently playing audio, the currently playing audio is matched with the reference rhythm information to obtain a first matching result, specifically including: When the rhythm display mode is the active display mode, the currently playing audio from the cloud-edge-device collaborative system is determined, and the currently playing audio is aligned with the corresponding baseline rhythm information according to the timestamp alignment rules to obtain the first matching result; wherein, the first matching result includes at least a rhythm preview file after the rhythm and audio are aligned. When the rhythm display mode is the passive display mode, the currently playing audio from the external environment is obtained, the currently playing audio is recognized, and the corresponding pre-associated rhythm information is loaded from the cloud storage module or local storage module according to the audio recognition result to obtain the corresponding first matching result. The audio segments of the currently playing audio are extracted according to a preset time interval, and the audio segments are compared with the audio recognition result. Based on the comparison result, it is determined whether to reload the pre-associated rhythm information and update the first matching result. When the rhythm display mode is the mixed display mode, the rhythm information corresponding to the first matching result and the music rhythm display information when the audio is played simultaneously are obtained, and the music rhythm display information is matched with the pre-stored rhythm information. In the case that the second matching result is inconsistent, the rhythm point timestamp corresponding to the music rhythm display information is corrected based on the pre-stored rhythm information to dynamically calibrate the playback sequence of the multi-dimensional effect display module and update the first matching result.
4. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 2, characterized in that, The method further includes: When the operating mode is the rhythm generation mode, the motion posture displacement data is determined; wherein, the motion posture displacement data is acquired through one or more of the following methods: sensor acquisition, motion posture video analysis and generation, and manual control parameter setting; The first rhythm information is determined based on the motion posture displacement data; The first rhythm information is input into a preset heterogeneous rhythm effect mapping model to generate second rhythm information corresponding to each of the currently connected rhythm effect sub-modules. The preset heterogeneous rhythm effect mapping model generates a first rhythm vector corresponding to the first rhythm information and converts the rhythm vector into a second rhythm vector corresponding to the rhythm effect sub-module according to the module category of the rhythm effect sub-module. The first rhythm vector and each of the second rhythm vectors have a multi-dimensional association relationship. This multi-dimensional association relationship includes at least mapping relationship type, dimension adaptation rules, and parameter mapping logic. Add the first rhythm information and the second rhythm information to the rhythm information.
5. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 1, characterized in that, After generating rhythm information, the method further includes: The number of groups N and the grouping criteria are received in real time; the grouping criteria include at least time interval, rhythm point density or scale type. According to the grouping criteria, the rhythm points in the rhythm information are divided into N rhythm groups, and the display intensity coefficient of each rhythm group is determined so as to control the effect display of the rhythm information according to the corresponding display intensity coefficient; the display intensity coefficient corresponds to at least one or more of the following: light intensity, vibration amplitude, image transparency, and jet system intensity.
6. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 1, characterized in that, The delay compensation time is calculated based on the probe feedback information corresponding to the test data packet and the pre-trained delay compensation model, specifically including: Calculate the corresponding transmission delay based on the feedback information and the delay calculation formula corresponding to the corresponding preset transmission architecture; The transmission delay and preset delay impact parameters are input into a pre-trained delay compensation model to determine the delay compensation time; the preset delay impact parameters include at least transmission distance, signal strength and channel utilization.
7. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 1, characterized in that, The method further includes: Determine the audio features corresponding to the rhythm information; wherein, the audio features include at least time-domain features, frequency-domain features, harmony features, meter features, and timbre features; The audio features are input into a pre-trained music genre detection model to determine the corresponding music genre label; Based on the style tag and the preset rhythm group adjustment list, determine the style adjustment parameter group corresponding to the rhythm information, and adjust the control parameters of each rhythm effect sub-module according to the style adjustment parameter group so that the displayed rhythm effect conforms to the style tag.
8. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 1, characterized in that, The method further includes: When the operating mode is interactive perception mode, determine each rhythm point corresponding to the reference rhythm information; At a preset time before the timestamp of each rhythm point, a weak effect prompt instruction is generated so that the corresponding rhythm effect submodule displays the corresponding weak rhythm effect to remind the user to perform interactive operations. Based on the accuracy of the user's interactive operation triggered by the weak rhythm effect, the rhythm level corresponding to the baseline rhythm information is corrected; wherein, different rhythm levels correspond to rhythm information with different rhythm groups. The method further includes: When the operating mode is interactive perception mode, the interaction rhythm matching degree corresponding to the user's interactive operation is determined, so as to adjust the controllable rhythm variable according to the interaction rhythm matching degree and generate a rhythm effect corresponding to the controllable rhythm variable; the rhythm effect includes at least a rhythm expansion effect and a rhythm decay effect.
9. The method for multi-mode, multi-dimensional perception and effect generation and display of music rhythm according to claim 8, characterized in that, The method further includes: In response to a rhythm practice request from the human-computer interaction module, based on the trigger accuracy of the user's interactive operation according to the weak rhythm effect and a preset accuracy threshold, a corresponding preset rhythm discard mode or preset rhythm expansion mode is determined to update the baseline rhythm information; wherein, the preset rhythm discard mode is used to remove rhythm points based on discard rules, the discard rules including at least one of random, grouped, dense, sparse, speed matching, and AI discard; the preset rhythm expansion mode is used to add rhythm information or rhythm groups based on preset expansion rules.
10. A multi-mode, multi-dimensional perception and effect generation and display system for music rhythm, characterized in that, The system includes: The receiving module is used to receive mode selection instructions from the human-computer interaction module to determine the operating mode as rhythm generation mode or rhythm display mode; wherein, the rhythm generation mode includes manual generation mode, automatic generation mode and semi-automatic generation mode, and the rhythm display mode includes at least the following sub-modes: active display mode, passive display mode and hybrid display mode. The generation module is used to generate rhythm information based on user input and / or system automatic analysis results when the operating mode is rhythm generation mode; wherein, the rhythm information includes at least rhythm point timestamps, control parameters of the corresponding rhythm effect submodule, and effect mapping rules. The matching module is used to, when the running mode is rhythm display mode, take the pre-generated rhythm information as the reference rhythm information, and perform feature matching between the currently playing audio and the reference rhythm information based on the source information of the currently playing audio to obtain a first matching result. The sending calculation module is used to send test data packets to each rhythm effect submodule, and calculate the delay compensation time based on the detection feedback information corresponding to the test data packets and the pre-trained delay compensation model, so as to adjust the trigger time of each rhythm point in the rhythm information according to the delay compensation time. The sending module is used to send control commands to each of the rhythm effect sub-modules according to the adjusted trigger time, so as to drive each of the rhythm effect sub-modules to synchronously display the corresponding rhythm effect according to the control commands.