Intelligent music playing method based on three-dimensional scene interaction

By extracting music feature vectors in the smart music player and generating a three-dimensional scene, materializing the control components and driving the dynamic elements, the problem of insufficient three-dimensional interaction in the existing technology is solved, and a highly immersive and high-freedom three-dimensional interactive experience is achieved.

CN120803391AActive Publication Date: 2025-10-17RIVOTEK TECH (JIANGSU) CO LTD

Patent Information

Application Number
CN202510946661.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing smart music players lack efficient multi-dimensional feature mapping and user interaction support in three-dimensional scene interaction, resulting in insufficient immersion and control experience, and failing to achieve spatiotemporal consistency and high degree of freedom in the interaction between music rhythm and visual elements.

Method used

By extracting music feature vectors, a lightweight neural network is used to generate a 3D scene, materialize control components, detect the user's hand posture, and drive dynamic elements such as light beam amplitude and particle aggregation to achieve a natural and intuitive 3D interactive experience.

Benefits of technology

It improves system response efficiency and user control smoothness, enhances visual expression and emotional resonance, and achieves highly immersive and consistent intelligent music playback effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803391A_ABST
    Figure CN120803391A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent music playing method based on three-dimensional scene interaction, and relates to the technical field of three-dimensional interaction, and the method comprises the steps: extracting a music feature vector of an input audio, inputting a lightweight neural network to output a music genre label, loading a three-dimensional scene in a three-dimensional scene template library based on the music genre label, and generating an initial three-dimensional music space; materializing a control component in the three-dimensional scene into a three-dimensional interactive object, and detecting the position and posture of the hand of the user in the initial three-dimensional music space to perform music playing control; and driving a dynamic element in the three-dimensional scene based on the music feature vector to control the amplitude of a top light beam, the ripple radius and the particle spacing of the central region, and mapping to the initial three-dimensional music space. According to the method, playing, volume and lyric functions are materialized into interactive objects, gesture recognition is introduced, and natural, visual and low-delay three-dimensional interactive experience is achieved; and an intelligent music playing effect with high immersion, high degree of freedom and high consistency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional interaction, and particularly to an intelligent music playing method based on three-dimensional scene interaction. BACKGROUND

[0002] With the rapid evolution of computer graphics, virtual reality (VR) and human-computer interaction (HCI) technologies, interactive applications based on three-dimensional scenes have attracted widespread attention in the fields of entertainment, education, industrial simulation, etc. In recent years, with the increasing feasibility of deploying deep learning and lightweight neural network models on mobile terminals and embedded devices, audio analysis and semantic recognition technologies have made significant progress. Among them, the multi-modal feature extraction method based on short-time Fourier transform (STFT), convolutional neural network (CNN) and recurrent neural network (RNN) can obtain key parameters such as beat per minute (BPM), spectral centroid, loudness and spectral energy distribution in real time, providing reliable data support for music genre classification and emotion analysis. At the same time, the functions of three-dimensional rendering engines (such as Unreal Engine, Unity) in lighting models, particle systems and physical simulations are increasingly perfect, making it possible to balance between smoothness and visual effects. In addition, adaptive scene control, gesture recognition and pose estimation based on deep reinforcement learning (DRL) are also gradually mature, laying a solid foundation for building immersive and interactive three-dimensional music spaces.

[0003] However, the existing intelligent music players mostly stay in two-dimensional interface interaction or simple visual synchronization level, failing to fully utilize the immersive experience brought by three-dimensional scenes and physical spaces. On the one hand, the existing solutions lack efficient connection between music genre classification and visual driving, often rendering with predefined animations or fixed templates, making it difficult to realize dynamic multi-dimensional feature mapping; on the other hand, most three-dimensional music visualization systems have limited support for user interaction, only providing two-dimensional operation methods such as clicking or mouse dragging, which cannot intuitively capture the influence of hand position and posture on playback control, limiting the natural integration of users and scenes. In addition, there are few attempts to deeply bind the particle system driven by audio features and scene entity components, resulting in insufficient spatio-temporal consistency between music rhythm, frequency band and visual elements (such as light column amplitude, ripple diffusion, particle aggregation), which cannot realize fine linkage and interactive feedback of music feature vectors in three-dimensional space. These deficiencies directly affect the user's sense of immersion and control experience, and also make it difficult to meet the professional needs of high degree of freedom and low latency interaction. SUMMARY

[0004] In view of the problems existing in the prior art intelligent music playing method based on three-dimensional scene interaction, the present application is proposed. Therefore, the problem to be solved by the present application is how to provide an intelligent music playing method based on three-dimensional scene interaction.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In the first aspect, the present application provides an intelligent music playing method based on three-dimensional scene interaction, which includes extracting a music feature vector of an input audio, inputting the music feature vector into a lightweight neural network to output a music genre label, loading a three-dimensional scene in a three-dimensional scene template library based on the music genre label, and generating an initial three-dimensional music space;

[0007] Performing scene entity component initialization, virtualizing control component entities in the three-dimensional scene into three-dimensional interactive objects, and detecting the position and posture of a user's hand in the initial three-dimensional music space to control music playing;

[0008] Based on the music feature vector, driving dynamic element control in the three-dimensional scene to control the amplitude of the top light column, the ripple radius, and the particle spacing in the central area, and mapping to the initial three-dimensional music space.

[0009] As a preferred scheme of the intelligent music playing method based on three-dimensional scene interaction, the music feature vector includes beat rate, spectral centroid, high-low frequency energy ratio, and loudness level.

[0010] As a preferred scheme of the intelligent music playing method based on three-dimensional scene interaction, the extracting of the music feature vector of the input audio includes:

[0011] Extracting the music feature vector for each frame of audio, including beat rate, spectral centroid, high-low frequency energy ratio, and loudness level;

[0012] Calculating the short-time energy envelope for each frame of audio, obtaining the beat interval sequence, and obtaining the beat rate;

[0013] Performing short-time Fourier transform on each frame of audio to obtain the amplitude spectrum, and obtaining the spectral centroid and high-low frequency energy ratio;

[0014] Based on the root mean square energy of each frame of audio, converting to generate a real-time loudness level.

[0015] As a preferred scheme of the intelligent music playing method based on three-dimensional scene interaction, the virtualizing of the control component entities in the three-dimensional scene into three-dimensional interactive objects includes:

[0016] In the initial three-dimensional music space, reserving corresponding spatial positions for playing / pausing, volume adjustment, and lyrics display, and creating three-dimensional placeholder objects;

[0017] The virtual phonograph turntable is located in the front center of the scene, the light column array control is distributed on both sides of the turntable or around the base, arranged at equal intervals, and the lyrics particle rendering area is located in the upper center of the field of view.

[0018] As a preferred scheme of the intelligent music playing method based on three-dimensional scene interaction, wherein: the detection of the position and posture of the user's hand in the initial three-dimensional music space for music playing control comprises:

[0019] Continuously detecting the position and posture of the user's hand in the initial three-dimensional music space;

[0020] When the user's hand enters the virtual gramophone turntable operation radius and makes a grabbing posture, it is determined that the turntable control starts;

[0021] Mapping the playing progress and the turntable rotation, calculating the turntable rotation angle according to the current playing progress;

[0022] When the user's hand gesture is detected along the disc surface around the center axis, the system measures the change of the turntable rotation angle driven by the gesture in real time, and updates the playing progress in reverse;

[0023] After the system receives the new playing progress ratio, it immediately adjusts the audio playing position and synchronously updates the visual feedback;

[0024] When the user's hand touches the top of the light column and makes an up and down sliding gesture, it is determined that the volume adjustment interaction is triggered, and the volume control and light column height mapping are performed, and the height of each light column is calculated according to the loudness level;

[0025] When the user makes an up and down sliding gesture, the system measures the movement distance of the gesture in the vertical direction, maps it to the volume increase or decrease, and updates the height of all light columns;

[0026] The flick or tap action in the lyrics area triggers the lyrics display mode switching or annotation jumping, and the current lyrics segment and the voice intelligibility are extracted from the audio in real time, and the particle aggregation density is adjusted according to the voice intelligibility;

[0027] The lyrics content is mapped to a particle group, which is arranged in order in a translucent plane or a free space, and the particle spacing and cluster size are determined by the particle aggregation density.

[0028] As a preferred scheme of the intelligent music playing method based on three-dimensional scene interaction, wherein: the expression of the turntable rotation angle is:

[0029] θ = θ min + (θ max - θ min ) × P

[0030] Where: θ is the rotation angle of the turntable, θ min is the minimum angle of the turntable rotation, θ max is the maximum angle of the turntable rotation, and P is the playing progress ratio;

[0031] The expression of the height of each light column is:

[0032] H=k*L

[0033] Wherein: H is the height of each light column, k is the device adaptive calibration coefficient, which is pre-set according to scene size and visual effect; L is the loudness level;

[0034] The expression of the particle aggregation density is:

[0035] ρ=ρ min +(ρ max -ρ min )xC v

[0036] Wherein: ρ is the particle aggregation density, ρ min is the preset most sparse particle aggregation density, ρ max is the preset most dense particle aggregation density, C v is the voice intelligibility.

[0037] As a preferred scheme of the intelligent music playing method based on three-dimensional scene interaction, wherein: the music feature vector is used to drive the dynamic element in the three-dimensional scene to control the top light column amplitude, the ripple radius and the center area particle spacing, including:

[0038] The latest music feature vector and voice intelligibility are received, and a visualization update loop is triggered to obtain the audio high-frequency energy of the current frame to control the top light column amplitude, to make periodic up and down displacement of each top light column in the vertical direction, and the frequency is synchronized with the beat rate; the low-frequency energy of the current frame is obtained to control the ripple radius of the newly added ripple of each frame; and the voice intelligibility is used to control the center area particle spacing;

[0039] The instantaneous amplitude of the light column, the ripple radius and the center area particle spacing are respectively mapped to the three-dimensional space.

[0040] In a second aspect, the present application provides a computer device, including a memory and a processor, the memory stores a computer program, wherein: the processor implements the steps of the intelligent music playing method based on three-dimensional scene interaction when executing the computer program.

[0041] In a third aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein: the computer program is executed by a processor to implement the steps of the intelligent music playing method based on three-dimensional scene interaction.

[0042] The application has the beneficial effects that: the application realizes efficient mapping of audio content to a three-dimensional visual scene through deep coupling of a multi-dimensional music feature vector and a lightweight neural network; natural, intuitive and low-latency three-dimensional interactive experience is realized by virtualizing the playing, volume and lyrics functions into interactive objects and introducing gesture recognition; immersive visual feedback highly synchronized with music rhythm and emotion is realized by accurately mapping rhythm, high and low frequency and vocal clarity to light column amplitude, ripple diffusion and particle aggregation through multi-dimensional dynamic element driving. The system response efficiency and user control fluency are improved, the visual expressiveness and emotional resonance are enhanced, and the intelligent music playing effect with high immersion, high degree of freedom and high consistency is achieved. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0044] Figure 1 The flowchart of the intelligent music playing method based on three-dimensional scene interaction. DETAILED DESCRIPTION

[0045] In order to make the above-mentioned purposes, features and advantages of the present application more apparent, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0046] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from the description, and those skilled in the art can make similar generalizations without departing from the concept of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0047] Secondly, the term "one embodiment" or "embodiments" herein refers to a specific feature, structure or property that can be included in at least one implementation of the present application. The term "one embodiment" appearing in different places in the specification does not refer to the same embodiment, nor does it refer to an embodiment that is separate or selectively excluded from other embodiments.

[0048] REFERENCE Figure 1 For the first embodiment of the present application, the embodiment provides an intelligent music playing method based on three-dimensional scene interaction, comprising:

[0049] S1: extract the music feature vector of the input audio, input the music feature vector into a lightweight neural network to output a music genre label, load a three-dimensional scene in a three-dimensional scene template library based on the music genre label, and generate an initial three-dimensional music space;

[0050] Specifically, the music feature vector (including rhythm feature value, spectral distribution, loudness level) of the input audio is extracted in real time through the audio analysis module, and a preset three-dimensional scene template library is loaded.

[0051] Based on the music genre identifier in the music feature vector, the corresponding three-dimensional scene parameter set (including building model, basic lighting tone, particle motion algorithm) is dynamically activated to generate a rotatable / zoomable initial three-dimensional music space.

[0052] The system continuously collects raw audio data from input devices (such as microphones or audio interfaces) at a fixed sampling rate and data block length for framing;

[0053] The music feature vector is extracted for each frame of audio, including beat rate, spectral centroid, high-low frequency energy ratio, and loudness level.

[0054] The short-time energy envelope is calculated for each frame of audio to obtain the beat interval sequence and obtain the beat rate. The short-time Fourier transform is performed on each frame of audio to obtain the amplitude spectrum, and the spectral centroid and high-low frequency energy ratio are further calculated to reflect the frequency band distribution of the music as a whole. Based on the root mean square (RMS) energy of each frame of audio, the real-time loudness level is converted and generated, which is used for subsequent mapping of lighting intensity and interactive controls.

[0055] The music feature vector is input into a pre-trained classification model (such as a lightweight neural network or decision tree) to output a music genre label, such as electronic, classical, rock, etc.

[0056] Load the three-dimensional scene template library, read the template metadata stored locally or remotely at startup, including the corresponding building model, default lighting, and particle motion algorithm for each genre.

[0057] Whenever a new genre label is detected, the system looks up the corresponding three-dimensional scene template through an internal mapping table and calls the asynchronous interface of the three-dimensional engine to start loading the building model and particle system, while applying the lighting parameters to the scene environment.

[0058] After completing the loading of the building model, lighting, and particle system, a music scene base that can be freely rotated and scaled in a three-dimensional perspective is constructed as an initial three-dimensional music space.

[0059] S2: Perform scene entity component initialization, virtualize control component entities in the three-dimensional scene into three-dimensional interactive objects, and detect the position and posture of the user's hand in the initial three-dimensional music space for music playback control;

[0060] Specifically, perform scene entity component initialization, after the three-dimensional music scene base is ready, reserve corresponding spatial positions for play / pause, volume adjustment, and lyrics display, and create three-dimensional placeholder objects: a virtual phonograph turntable is located in the center of the scene in front, with a radius matching the user's comfortable operation range; the light column array control is distributed on both sides of the turntable or around the base, arranged at equal intervals; the lyrics particle rendering area is located in the center of the field of view above, forming a semi-transparent subtitle strip effect.

[0061] Start gesture recognition and continuously detect the position and posture of the user's hand in the initial three-dimensional music space; when the user's hand enters the virtual phonograph turntable operation radius and makes a grabbing gesture (such as five fingers open and then clenched), it is determined that the turntable control starts;

[0062] Perform playback progress and turntable rotation mapping, define the linear relationship between the current playback progress (range 0, 1) and the turntable rotation angle (unit: degree), represented as:

[0063] θ = θ min + (θ max - θ min ) × P

[0064] Where: θ is the rotation angle of the turntable, θ min is the minimum angle of the turntable rotation, θ max is the maximum angle of the turntable rotation, and P is the playback progress ratio (0 indicates start, 1 indicates end).

[0065] When a user rotates the hand gesture along the disc surface around the center axis is detected, the change of the rotation angle of the turntable driven by the gesture is measured in real time, and the playback progress is updated in reverse, represented as:

[0066]

[0067] After the system receives the new playback progress ratio, it immediately adjusts the audio playback position and synchronously updates the visual feedback.

[0068] When the user's hand touches the top of the light column and makes an up and down sliding gesture, it is determined as a volume adjustment interaction;

[0069] Perform volume control and light column height mapping to establish a linear relationship between the height of each light column and the loudness level (unit: dB or normalized value), represented as:

[0070] H = k × L

[0071] Wherein: H is the height of each light column, k is the device adaptive calibration coefficient, which is pre-set according to the scene size and visual effect; L is the loudness level.

[0072] When the user makes up and down sliding gestures, the system measures the movement distance of the gesture in the vertical direction, maps it to the volume increase and decrease, and updates the height of all light columns to reflect the latest loudness level, which is represented as:

[0073] ΔL = α × Δy

[0074] Wherein: ΔL is the change of loudness level, α is the mapping coefficient of spatial coordinates to loudness change, and Δy is the movement distance of the gesture in the vertical direction;

[0075] In the lyrics area, a swipe or tap action can trigger lyrics display mode switching or annotation jumping. Real-time extraction of the current lyrics segment and vocal clarity from the audio, dynamic adjustment of particle aggregation density according to vocal clarity, represented as:

[0076] ρ = ρ min + (ρ max - ρ min ) × C v

[0077] Wherein: ρ is the particle aggregation density, ρ min is the preset sparse particle aggregation density, ρ max is the preset dense particle aggregation density, and C v is the vocal clarity;

[0078] Map the lyrics content to the particle group, arrange it in order in the translucent plane or free space, and determine the particle spacing and cluster size with the particle aggregation density to form a three-dimensional lyrics effect that can be roamed and scaled.

[0079] Whenever the user completes an interaction (such as releasing the turntable, sliding ends), the system records the new playback progress ratio, loudness level, and particle aggregation density and modifies the synchronization to the audio engine and rendering engine;

[0080] After the interaction ends, if the user is stationary for more than a preset duration, the control controls are automatically hidden, only the dynamic visualization elements are displayed, and the overall picture beauty is maintained.

[0081] After completing the control binding and basic interaction process, the system will package the real-time playback progress ratio, loudness level, vocal clarity, and music feature vector collected, and pass it to the multi-dimensional visualization feedback module to drive the dynamic scene elements in the next stage.

[0082] S3: Based on the music feature vector, drive the dynamic elements in the three-dimensional scene to control the top light column amplitude, ripple radius, and central area particle spacing, which are mapped to the initial three-dimensional music space.

[0083] Specifically, based on the real-time updated music feature vector, drive the dynamic elements in the three-dimensional scene: high-frequency spectrum value control the amplitude of the top light column, low-frequency spectrum value generates ground wave diffusion effect, vocal feature value adjusts the central particle aggregation degree;

[0084] The system receives the latest music feature vector and vocal clarity in units of frames (e.g. every 50ms a frame) and triggers the visualization update loop.

[0085] Get the audio high-frequency energy of the current frame to control the amplitude of the top light column, and make periodic up-down displacement of each top light column in the vertical direction, the amplitude of which is dynamically adjusted according to the instantaneous amplitude of the light column, and the frequency is synchronized with the beat rate. Get the low-frequency energy of the current frame to control the wavelet radius of each new wavelet. Control the center particle spacing in the center area by the vocal clarity to adjust the degree of crowding of the central particle group. The greater the vocal clarity, the smaller the center area particle spacing, and the more closely the particles are gathered, forming a vocal core visual focus point.

[0086] Map the instantaneous amplitude of the light column, the wavelet radius and the center area particle spacing to the three-dimensional space respectively: the light column group is distributed on the top plate plane according to the fixed grid, each grid center hangs a light column, and the instantaneous amplitude of the light column corresponds to the vertical offset. The ground wave is drawn as a concentric ring on the floor plane with the scene coordinate system origin as the center. The center particles are randomly distributed in the spherical or planar area around the origin, and the current position of each particle needs to ensure that the distance between any two adjacent particles is greater than the center area particle spacing.

[0087] The embodiment also provides a computer device suitable for the intelligent music playing method based on three-dimensional scene interaction, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize all or part of the steps of the method described in the above embodiment.

[0088] The embodiment further provides a storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the method in any optional implementation manner of the above-mentioned embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk, or an optical disk.

[0089] The storage medium proposed in the embodiment belongs to the same inventive concept as the data storage method proposed in the above-mentioned embodiments, and the technical details not described in the embodiment can be referred to the above-mentioned embodiments, and the embodiment has the same beneficial effects as the above-mentioned embodiments.

[0090] To sum up, the application realizes efficient mapping of audio content to a three-dimensional visual scene through deep coupling of multi-dimensional music feature vectors and lightweight neural networks; realizes natural, intuitive, and low-latency three-dimensional interactive experience by virtualizing the play, volume, and lyrics functions into interactive objects and introducing gesture recognition; realizes immersive visual feedback highly synchronized with music rhythm, emotion, and clarity of vocals by accurately mapping rhythm, high and low frequencies, and vocal clarity to light column amplitude, ripple diffusion, and particle aggregation. This not only improves system response efficiency and user control smoothness, but also enhances visual expressiveness and emotional resonance, achieving high immersion, high degree of freedom, and high consistency in intelligent music playback.

[0091] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and all of them should be covered in the scope of the claims of the present application.

Claims

1. An intelligent music playing method based on three-dimensional scene interaction, characterized by: include, Extract the music feature vector of the input audio, input the music feature vector into a lightweight neural network to output a music genre label, and load the 3D scene from the 3D scene template library based on the music genre label to generate an initial 3D music space; Initialize the scene entity components, materialize the control components in the 3D scene into 3D interactive objects, and detect the position and posture of the user's hand in the initial 3D music space to control music playback; Based on the music feature vector, the dynamic elements in the three-dimensional scene are driven to control the amplitude of the top light column, the ripple radius and the particle spacing in the central area, and mapped into the initial three-dimensional music space.

2. The intelligent music playing method based on three-dimensional scene interaction according to claim 1, characterized in that: The music feature vector includes beat rate, spectral centroid, high-frequency and low-frequency energy ratio, and loudness level.

3. The intelligent music playing method based on three-dimensional scene interaction according to claim 2, characterized in that: Extracting the music feature vector of the input audio comprises: Extract music feature vectors for each audio frame, including beat rate, spectral centroid, high-low frequency energy ratio, and loudness level; Calculate the short-time energy envelope of each frame of audio to obtain the beat interval sequence and the beat rate; Perform short-time Fourier transform on each frame of audio to obtain the amplitude spectrum, spectral centroid and high-frequency and low-frequency energy ratio; The transform generates a real-time loudness level based on the RMS energy of each frame of audio.

4. The intelligent music playing method based on three-dimensional scene interaction according to claim 3, characterized in that: The step of materializing the control components in the three-dimensional scene into three-dimensional interactive objects includes: In the initial 3D music space, reserve corresponding spatial locations for play / pause, volume adjustment, and lyrics display, and create 3D placeholder objects; The virtual phonograph turntable is located in the front center of the scene. The light column array controls are distributed on both sides of the turntable or around the base, arranged at equal intervals. The lyrics particle rendering area is located above the center of the field of view.

5. The intelligent music playing method based on three-dimensional scene interaction according to claim 4, characterized in that: The detecting the position and posture of the user's hand in the initial three-dimensional music space to control the music playback includes: Continuously detect the position and posture of the user's hands in the initial three-dimensional music space; When the user's hand enters the operating radius of the virtual phonograph turntable and makes a grabbing gesture, it is determined that the turntable control has started; Map the playback progress to the turntable rotation, and calculate the turntable rotation angle based on the current playback progress; When a user's rotation gesture around the central axis of the turntable is detected, the turntable's rotation angle caused by the gesture is measured in real time, and the playback progress is updated in reverse. After the system receives the new playback progress ratio, it immediately adjusts the audio playback position and updates the visual feedback synchronously; When the user touches the top of the light column and swipes up or down, it is considered a volume adjustment interaction. The volume control is mapped to the light column height, and the height of each light column is calculated based on the loudness level. When the user swipes up or down, the system measures the vertical distance of the gesture, maps it to volume increase or decrease, and updates the height of all light beams. A swipe or tap in the lyrics area triggers a switch in the lyrics display mode or an annotation jump. The system extracts the current lyrics segment and vocal clarity from the audio in real time, and adjusts the particle aggregation density based on the vocal clarity. The lyrics content is mapped into a particle group, arranged in order on a translucent plane or in free space, and the particle spacing and cluster size are determined by the particle aggregation density.

6. The intelligent music playing method based on three-dimensional scene interaction according to claim 5, characterized in that: The expression of the turntable rotation angle is: θ=θ min +(θ max -θ min )×P Where: θ is the rotation angle of the turntable, θ min is the minimum angle that the turntable can rotate, θ max is the maximum angle that the turntable can rotate, and P is the playback progress ratio; The expression of the height of each light beam is: H=k×L Where: H is the height of each light beam, k is the device adaptive calibration coefficient, which is pre-set according to the scene size and visual effect; L is the loudness level; The expression of the particle aggregation density is: p=p min +(r max -r min )×C v Where: ρ is the particle aggregation density, ρ min is the preset sparsest particle aggregation density, ρ max is the preset most dense particle aggregation density, C v For vocal clarity.

7. The intelligent music playing method based on three-dimensional scene interaction according to claim 6, characterized in that: The method of driving the dynamic elements in the three-dimensional scene based on the music feature vector to control the top light column amplitude, ripple radius and center area particle spacing includes: Receive the latest music feature vector and vocal clarity, and trigger the visualization update loop. The high-frequency energy of the current frame's audio is used to control the amplitude of the top light beams. Each top light beam is periodically shifted vertically, with the frequency synchronized with the beat rate. The low-frequency energy of the current frame is used to control the ripple radius of each new ripple. The vocal clarity is used to control the spacing between particles in the central area. The instantaneous amplitude of the light column, the ripple radius and the distance between particles in the central area are mapped into three-dimensional space respectively.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent music playback method based on three-dimensional scene interaction according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent music playback method based on three-dimensional scene interaction according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Interactive music visualization method and device

    CN104732983A

  • Audio playing control method and device, equipment and storage medium

    CN120215870A

  • Information displaying method and terminal

    US20190215397A1

Cited By

  • Real-time music generation and sound and picture linkage method and system based on voice feature analysis

    CN122290628A

  • A Real-Time Music Generation and Audio-Visual Synchronization Method and System Based on Speech Feature Analysis

    CN122290628B