Intelligent music playing method based on three-dimensional scene interaction

By extracting music feature vectors from a smart music player and materializing control components, and combining this with gesture recognition to drive dynamic elements, the problem of insufficient immersion and control experience in 3D scene interaction is solved, achieving a highly immersive and highly free 3D interactive experience.

CN120803391BActive Publication Date: 2026-03-17RIVOTEK TECH (JIANGSU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing smart music players lack efficient multi-dimensional feature mapping and user interaction support in three-dimensional scene interaction, resulting in insufficient immersion and control experience, and failing to achieve spatiotemporal consistency and high degree of freedom in the interaction between music rhythm and visual elements.

Method used

By extracting music feature vectors and inputting them into a lightweight neural network, an initial three-dimensional music space is generated. The control components are materialized and the user's hand posture is detected to drive dynamic elements such as light beam amplitude and particle aggregation, achieving a natural and intuitive three-dimensional interactive experience.

Benefits of technology

It achieves precise linkage and highly immersive interaction of music feature vectors in three-dimensional space, improves system response efficiency and user operation smoothness, and enhances visual expressiveness and emotional resonance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803391B_ABST
    Figure CN120803391B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent music playing method based on three-dimensional scene interaction and relates to the technical field of three-dimensional interaction, which comprises the following steps: extracting a music feature vector of input audio, inputting the music feature vector into a lightweight neural network to output a music genre label, loading a three-dimensional scene in a three-dimensional scene template library based on the music genre label to generate an initial three-dimensional music space, and solidifying a control component in the three-dimensional scene into a three-dimensional interactive object and detecting the position and posture of a user's hand in the initial three-dimensional music space to control music playing; and driving dynamic elements in the three-dimensional scene based on the music feature vector to control the amplitude of a top light column, the radius of a ripple and the particle spacing in a central area, and mapping the three-dimensional scene to the initial three-dimensional music space. The application solidifies playing, volume and lyrics functions into interactive objects and introduces gesture recognition, so that natural, intuitive and low-delay three-dimensional interactive experience is realized, and the application achieves the effects of high immersion, high freedom and high consistency in intelligent music playing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional interactive technology, and in particular to an intelligent music playback method based on three-dimensional scene interaction. Background Technology

[0002] With the rapid evolution of computer graphics, virtual reality (VR), and human-computer interaction (HCI) technologies, interactive applications based on 3D scenes have gained widespread attention in entertainment, education, and industrial simulation. In recent years, with the increasing feasibility of deploying deep learning and lightweight neural network models on mobile and embedded devices, audio analysis and semantic recognition technologies have made significant progress. Among these advancements, multimodal feature extraction methods based on Short-Time Fourier Transform (STFT), Convolutional Neural Networks (CNN), and Recurrent Neural Networks (RNN) can acquire key parameters such as beat per minute (BPM), spectral centroid, loudness, and spectral energy distribution in real time, providing reliable data support for music genre classification and sentiment analysis. Simultaneously, 3D rendering engines (such as Unreal Engine and Unity) have become increasingly sophisticated in lighting models, particle systems, and physics simulations, making it possible to achieve a balance between smoothness and visual effects. Furthermore, adaptive scene control, gesture recognition, and pose estimation based on Deep Reinforcement Learning (DRL) have matured, laying a solid foundation for constructing immersive and interactive 3D music spaces.

[0003] However, current smart music players mostly remain at the level of two-dimensional interface interaction or simple visualization synchronization, failing to fully utilize the immersive experience brought by three-dimensional scenes and physical space. On the one hand, existing solutions lack efficient connection between music genre classification and visualization-driven processes, often using predefined animations or fixed templates for rendering, making it difficult to achieve dynamic multi-dimensional feature mapping. On the other hand, most 3D music visualization systems offer limited support for user interaction, providing only two-dimensional operations such as clicking or dragging, failing to intuitively capture the impact of hand position and posture on playback control, thus limiting the natural integration of the user with the scene. In addition, there are few attempts to deeply bind audio feature-driven particle systems with scene entity components, resulting in insufficient spatiotemporal consistency between music rhythm, frequency bands, and visual elements (such as light beam amplitude, ripple diffusion, and particle aggregation), making it impossible to achieve fine linkage and interactive feedback as the music feature vector moves in three-dimensional space. These shortcomings directly affect the user's immersion and control experience, and also fail to meet the professional needs for high-degree-of-freedom, low-latency interaction. Summary of the Invention

[0004] In view of the problems existing in existing intelligent music playback methods based on 3D scene interaction, this invention is proposed. Therefore, the problem to be solved by this invention is how to provide an intelligent music playback method based on 3D scene interaction.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0006] In a first aspect, the present invention provides an intelligent music playback method based on three-dimensional scene interaction, which includes: extracting music feature vectors of input audio, inputting the music feature vectors into a lightweight neural network to output music genre labels, loading a three-dimensional scene from a three-dimensional scene template library based on the music genre labels, and generating an initial three-dimensional music space.

[0007] Initialize scene entity components, materialize control components in the 3D scene into 3D interactive objects, and detect the position and posture of the user's hand in the initial 3D music space to control music playback;

[0008] Based on the music feature vector, the dynamic elements in the 3D scene are controlled to control the amplitude of the top light column, the radius of the ripples, and the particle spacing in the central region, which are then mapped to the initial 3D music space.

[0009] As a preferred embodiment of the intelligent music playback method based on three-dimensional scene interaction described in this invention, the music feature vector includes beat rate, spectral centroid, high-low frequency energy ratio, and loudness level.

[0010] As a preferred embodiment of the intelligent music playback method based on three-dimensional scene interaction described in this invention, the extraction of the music feature vector of the input audio includes:

[0011] Extract music feature vectors from each frame of audio, including beat rate, spectral centroid, high-low frequency energy ratio, and loudness level;

[0012] Calculate the short-time energy envelope for each frame of audio to obtain the beat interval sequence and obtain the beat rate;

[0013] Perform a short-time Fourier transform on each frame of audio to obtain the amplitude spectrum, and obtain the spectral centroid and the high-low frequency energy ratio;

[0014] The real-time loudness level is generated based on the root mean square energy of each audio frame.

[0015] As a preferred embodiment of the intelligent music playback method based on three-dimensional scene interaction described in this invention, wherein: the step of materializing the control components in the three-dimensional scene into three-dimensional interactive objects includes:

[0016] In the initial three-dimensional music space, corresponding space positions are reserved for play / pause, volume adjustment and lyrics display, and three-dimensional placeholder objects are created;

[0017] The virtual phonograph turntable is located in the center of the scene, with light beam array controls distributed on both sides of the turntable or around the base, arranged at equal intervals. The lyrics particle rendering area is located above the center of the field of view.

[0018] As a preferred embodiment of the intelligent music playback method based on three-dimensional scene interaction described in this invention, the step of detecting the position and posture of the user's hand in the initial three-dimensional music space for music playback control includes:

[0019] Continuously detect the position and posture of the user's hands in the initial three-dimensional music space;

[0020] When the user's hand enters the operating radius of the virtual phonograph turntable and makes a grasping gesture, the turntable control is determined to begin.

[0021] Map the playback progress to the turntable rotation, and calculate the turntable rotation angle based on the current playback progress;

[0022] When a user's gesture of rotating the disc around the central axis is detected, the change in the rotation angle of the disc caused by the gesture is measured in real time, and the playback progress is updated accordingly.

[0023] Upon receiving the new playback progress ratio, the system immediately adjusts the audio playback position and updates the visual feedback simultaneously.

[0024] When a user touches the top of the light pillar and makes a swipe gesture, it is determined to be a volume adjustment interaction. The volume control is mapped to the height of the light pillar, and the height of each light pillar is calculated based on the loudness level.

[0025] When a user makes a swipe gesture, the system measures the vertical distance the gesture moves and maps it to volume up or down, while also updating the height of all light pillars.

[0026] A swipe or tap action within the lyrics area triggers a switch in the lyrics display mode or a jump to annotations. The current lyrics segment and vocal clarity are extracted from the audio in real time, and the particle aggregation density is adjusted according to the vocal clarity.

[0027] The lyrics are mapped as a swarm of particles, arranged sequentially in a semi-transparent plane or free space, with the particle spacing and cluster size determined by the particle aggregation density.

[0028] In a preferred embodiment of the intelligent music playback method based on three-dimensional scene interaction described in this invention, the expression for the rotation angle of the turntable is:

[0029] θ=θ min +(θ max -θ min )×P

[0030] Where: θ is the rotation angle of the turntable, θ min θ is the minimum angle at which the turntable can rotate. max P represents the maximum angle the turntable can rotate, and P is the playback progress percentage.

[0031] The expression for the height of each light beam is:

[0032] H = k × L

[0033] Where: H is the height of each light beam, k is the device adaptive calibration coefficient, which is preset according to the scene size and visual effect; L is the loudness level;

[0034] The expression for the particle aggregation density is:

[0035] ρ=ρ min +(ρ max -ρ min )×C v

[0036] Where: ρ is the particle aggregation density, ρ min ρ is the preset minimum sparse particle aggregation density. max C is the preset maximum particle aggregation density. v Voice clarity.

[0037] As a preferred embodiment of the intelligent music playback method based on three-dimensional scene interaction described in this invention, wherein: the control of the amplitude of the top light column, the radius of the ripples, and the particle spacing in the central region by driving dynamic elements within the three-dimensional scene based on music feature vectors includes:

[0038] Receive the latest music feature vector and vocal intelligibility, trigger the visualization update loop, obtain the high-frequency energy of the audio in the current frame to control the amplitude of the top light pillar, and periodically shift each top light pillar vertically, with the frequency synchronized with the beat rate; obtain the low-frequency energy of the current frame to control the ripple radius of the newly added ripples in each frame; and control the particle spacing in the central region through vocal intelligibility.

[0039] The instantaneous amplitude, ripple radius, and particle spacing in the central region of the light beam are mapped to three-dimensional space.

[0040] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of an intelligent music playback method based on three-dimensional scene interaction.

[0041] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements the steps of an intelligent music playback method based on three-dimensional scene interaction.

[0042] The beneficial effects of this invention are as follows: By deeply coupling multi-dimensional music feature vectors with lightweight neural networks, this invention achieves efficient mapping of audio content to a three-dimensional visual scene; by materializing playback, volume, and lyrics functions into interactive objects and introducing gesture recognition, it achieves a natural, intuitive, and low-latency three-dimensional interactive experience; through multi-dimensional dynamic element driving, rhythm, high and low frequencies, and vocal clarity are precisely mapped to light beam amplitude, ripple diffusion, and particle aggregation, achieving immersive visual feedback highly synchronized with musical beats and emotions. This not only improves system response efficiency and user control smoothness but also enhances visual expressiveness and emotional resonance, achieving a highly immersive, highly free, and highly consistent intelligent music playback effect. Attached Figure Description

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart of an intelligent music playback method based on 3D scene interaction. Detailed Implementation

[0045] To make the above-mentioned objects, features, and advantages of the present invention more readily understood, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0046] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0047] Secondly, the term "one embodiment" or "example" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. An embodiment appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment that selectively excludes other embodiments.

[0048] Reference Figure 1 This is the first embodiment of the present invention, which provides an intelligent music playback method based on three-dimensional scene interaction, including:

[0049] S1: Extract the music feature vector of the input audio, input the music feature vector into a lightweight neural network to output music genre labels, load the 3D scene in the 3D scene template library based on the music genre labels, and generate the initial 3D music space;

[0050] Specifically, the audio parsing module extracts the music feature vector (including rhythm feature value, spectrum distribution, and loudness level) of the input audio in real time, while loading a preset 3D scene template library.

[0051] Based on the music genre identifier in the music feature vector, the corresponding three-dimensional scene parameter set (including building model, basic lighting tone, particle motion algorithm) is dynamically activated to generate an initial three-dimensional music space that can be rotated / scaled.

[0052] The system continuously acquires raw audio data from input devices (such as microphones or audio interfaces) and divides it into frames with a fixed sampling rate and data block length;

[0053] Extract music feature vectors from each frame of audio, including beat rate, spectral centroid, high-low frequency energy ratio, and loudness level;

[0054] For each audio frame, a short-time energy envelope is calculated to obtain the beat interval sequence and thus the beat rate. A short-time Fourier transform is performed on each audio frame to obtain the amplitude spectrum. The spectral centroid and the high-low frequency energy ratio are then calculated to reflect the overall frequency distribution of the music. Based on the root mean square (RMS) energy of each audio frame, a real-time loudness level is generated for subsequent mapping of light intensity and interactive controls.

[0055] Input the music feature vector into a pre-trained classification model (such as a lightweight neural network or decision tree) and output music genre labels, such as electronic, classical, rock, etc.

[0056] Load the 3D scene template library and read the template metadata stored locally or remotely at startup, including the architectural model, default lighting, and particle motion algorithm corresponding to each genre.

[0057] Whenever a new genre tag is detected, the system looks up the corresponding 3D scene template through the internal mapping table, calls the asynchronous interface of the 3D engine to start loading the building model and particle system, and applies lighting parameters to the scene environment.

[0058] After completing the building model, lighting, and particle system loading, a music scene base that can be freely rotated and scaled in a three-dimensional view is constructed as the initial three-dimensional music space.

[0059] S2: Initialize scene entity components, materialize control components in the 3D scene into 3D interactive objects, and detect the position and posture of the user's hand in the initial 3D music space to control music playback.

[0060] Specifically, the scene entity components are initialized. After the 3D music scene base is ready, corresponding space positions are reserved for play / pause, volume adjustment and lyrics display, and 3D placeholder objects are created: the virtual phonograph turntable is located in the center front of the scene, with a radius that matches the user's comfortable operating range; the light beam array controls are distributed on both sides of the turntable or around the base, arranged at equal intervals; the lyrics particle rendering area is located above the center of the field of view, forming a semi-transparent subtitle effect.

[0061] Gesture recognition is activated to continuously detect the position and posture of the user's hand in the initial three-dimensional music space; when the user's hand enters the operating radius of the virtual phonograph turntable and makes a grasping gesture (such as spreading the five fingers and then making a fist), it is determined that turntable control has started.

[0062] Mapping playback progress to dial rotation, defining a linear relationship between the current playback progress (range 0, 1) and the dial rotation angle (unit: degrees), expressed as:

[0063] θ=θ min +(θ max -θ min )×P

[0064] Where: θ is the rotation angle of the turntable, θ min θ is the minimum angle at which the turntable can rotate. max P represents the maximum angle the turntable can rotate, and P represents the playback progress percentage (0 indicates start, 1 indicates end).

[0065] When a user's gesture of rotating the disc around its central axis is detected, the change in the rotation angle of the disc caused by the gesture is measured in real time, and the playback progress is updated accordingly, as shown below:

[0066]

[0067] Upon receiving the new playback progress ratio, the system immediately adjusts the audio playback position and updates the visual feedback simultaneously.

[0068] When a user touches the top of the light bar and makes a swipe gesture up or down, it is determined to be a volume adjustment interaction;

[0069] Volume control is mapped to the height of the light pillars, establishing a linear relationship between the height of each light pillar and the loudness level (unit: dB or normalized value), expressed as:

[0070] H = k × L

[0071] Where: H is the height of each light beam, k is the device adaptive calibration coefficient, which is preset according to the scene size and visual effect; L is the loudness level.

[0072] When a user makes a swipe gesture, the system measures the vertical distance the gesture moves, maps it to volume increase or decrease, and simultaneously updates the height of all light bars to reflect the latest loudness level, as shown below:

[0073] ΔL=α×Δy

[0074] Where: ΔL is the change in loudness level, α is the mapping coefficient from spatial coordinates to loudness change, and Δy is the vertical distance the gesture moves.

[0075] A swipe or tap within the lyrics area can trigger a switch in the lyrics display mode or a jump to annotations. The system extracts the current lyrics segment and vocal clarity from the audio in real time, dynamically adjusting the particle density based on vocal clarity, as shown below:

[0076] ρ=ρ min +(ρ max -ρ min )×C v

[0077] Where: ρ is the particle aggregation density, ρ min ρ is the preset minimum sparse particle aggregation density. max C is the preset maximum particle aggregation density. v Voice clarity;

[0078] The lyrics are mapped as a swarm of particles, arranged sequentially in a semi-transparent plane or free space, and the particle spacing and cluster size are determined by the particle aggregation density, forming a roamable and scalable three-dimensional lyrics effect.

[0079] Each time a user completes an interaction (such as releasing the turntable or ending a swipe), the system records the new playback progress percentage, loudness level, and particle aggregation density, and modifies and synchronizes these parameters to the audio engine and rendering engine.

[0080] After the interaction ends, if the user remains still for more than a preset time, the control controls will be automatically hidden, and only dynamic visual elements will be displayed to maintain the overall aesthetic appeal of the screen.

[0081] After completing the control binding and basic interaction process, the system will package the collected real-time playback progress ratio, loudness level, vocal clarity, and music feature vectors and pass them to the multi-dimensional visualization feedback module to drive the dynamic scene elements in the next stage.

[0082] S3: Based on the music feature vector, the dynamic elements in the 3D scene are controlled to control the amplitude of the top light column, the radius of the ripples, and the particle spacing in the central region, and mapped to the initial 3D music space.

[0083] Specifically, based on real-time updated music feature vectors, dynamic elements within the 3D scene are driven: high-frequency spectrum values ​​control the amplitude of the top light pillar, low-frequency spectrum values ​​generate the ground ripple diffusion effect, and human voice feature values ​​adjust the central particle aggregation degree.

[0084] The system receives the latest music feature vectors and vocal clarity in frames (e.g., one frame every 50ms) and triggers a visualization update loop.

[0085] The high-frequency energy of the audio in the current frame controls the amplitude of the top light pillars, periodically shifting each pillar vertically. The amplitude is dynamically adjusted according to the pillar's instantaneous amplitude, with the frequency synchronized with the beat rate. The low-frequency energy of the current frame controls the radius of the newly added ripples in each frame. The particle spacing in the central region is controlled by the voice clarity, adjusting the clustering of the central particle group. Higher voice clarity results in smaller particle spacing in the central region, causing the particles to cluster more tightly and forming a core visual focal point for the voice.

[0086] The instantaneous amplitude, ripple radius, and particle spacing in the central region of the light pillars are mapped to three-dimensional space: the light pillar group is distributed on the top plane according to a fixed grid, with one light pillar hanging at the center of each grid, and the instantaneous amplitude of the light pillars corresponding to the vertical offset. The ground ripples are drawn as concentric rings on the floor plane with the origin of the scene coordinate system as the center. The central particles are randomly distributed in a spherical or planar region around the origin, and the current position of each particle must ensure that the distance between any two adjacent particles is greater than the particle spacing in the central region.

[0087] This embodiment also provides a computer device applicable to the intelligent music playback method based on three-dimensional scene interaction, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement all or part of the steps of the method described in the above embodiments of the present invention.

[0088] This embodiment also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, it performs the method in any optional implementation of the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0089] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0090] In summary, this invention achieves efficient mapping of audio content to a three-dimensional visual scene through deep coupling of multi-dimensional music feature vectors and lightweight neural networks; by materializing playback, volume, and lyrics functions as interactive objects and introducing gesture recognition, it achieves a natural, intuitive, and low-latency three-dimensional interactive experience; and through multi-dimensional dynamic element driving, it accurately maps rhythm, high and low frequencies, and vocal clarity into light beam amplitude, ripple diffusion, and particle aggregation, achieving immersive visual feedback highly synchronized with music beats and emotions. This not only improves system response efficiency and user control smoothness but also enhances visual expressiveness and emotional resonance, achieving a highly immersive, highly flexible, and highly consistent intelligent music playback effect.

[0091] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An intelligent music playing method based on three-dimensional scene interaction, characterized in that: The method comprises the steps of: extracting a music feature vector of the input audio, inputting the music feature vector into a lightweight neural network to output a music genre label, loading a three-dimensional scene in a three-dimensional scene template library based on the music genre label, and generating an initial three-dimensional music space; initializing scene entity components, virtualizing control component entities in the three-dimensional scene into three-dimensional interactive objects, and detecting the position and posture of the user's hand in the initial three-dimensional music space to control music playing; driving dynamic elements in the three-dimensional scene based on the music feature vector to control the amplitude of the top light column, the ripple radius, and the particle spacing in the central area, and mapping to the initial three-dimensional music space; the driving dynamic elements in the three-dimensional scene based on the music feature vector to control the amplitude of the top light column, the ripple radius, and the particle spacing in the central area comprises: receiving the latest music feature vector and the vocal clarity, triggering a visualization update loop, obtaining the audio high-frequency energy of the current frame to control the amplitude of the top light column, periodically displacing each top light column in the vertical direction, and synchronizing the frequency with the beat rate; obtaining the low-frequency energy of the current frame to control the ripple radius of each newly added ripple; and controlling the particle spacing in the central area through the vocal clarity; mapping the instantaneous amplitude of the light column, the ripple radius, and the particle spacing in the central area to the three-dimensional space respectively. 2.The intelligent music playing method based on three-dimensional scene interaction of claim 1, wherein: The music feature vector comprises a beat rate, a spectral centroid, a high-low frequency energy ratio, and a loudness level. 3.The intelligent music playing method based on three-dimensional scene interaction of claim 2, wherein: The method for extracting the music feature vector of the input audio comprises: extracting a music feature vector for each frame of audio, including a beat rate, a spectral centroid, a high-low frequency energy ratio, and a loudness level; calculating a short-time energy envelope for each frame of audio, obtaining a beat interval sequence, and obtaining a beat rate; performing a short-time Fourier transform on each frame of audio to obtain an amplitude spectrum, and obtaining a spectral centroid and a high-low frequency energy ratio; based on the root mean square energy of each frame of audio, converting to generate a real-time loudness level. 4.The intelligent music playing method based on three-dimensional scene interaction of claim 3, wherein: The method for virtualizing control component entities in the three-dimensional scene into three-dimensional interactive objects comprises: reserving corresponding spatial positions for play / pause, volume adjustment, and lyric display in the initial three-dimensional music space, and creating three-dimensional placeholder objects; the virtual gramophone turntable is located in the center of the scene in front, the light column array control is distributed on both sides of the turntable or around the base, and is arranged at equal intervals, and the lyric particle rendering area is located in the upper center of the field of view. 5.The intelligent music playing method based on three-dimensional scene interaction of claim 4, wherein: The method for detecting the position and posture of the user's hand in the initial three-dimensional music space to control music playing comprises: continuously detecting the position and posture of the user's hand in the initial three-dimensional music space; when the user's hand enters the operation radius of the virtual gramophone turntable and makes a grabbing gesture, it is determined that the turntable control starts; mapping the play progress to the rotation of the turntable, and calculating the rotation angle of the turntable according to the current play progress; when a user gesture of rotating the hand along the disc surface around the central axis is detected, the rotation angle change of the turntable driven by the gesture is measured in real time, and the play progress is updated in reverse; after the system receives a new play progress ratio, the audio playing position is immediately adjusted and the visual feedback is updated synchronously; when the user's hand touches the top of the light column and makes an up-down sliding gesture, it is determined that the volume adjustment interaction is performed, the volume control is mapped to the height of the light column, and the height of each light column is calculated according to the loudness level. When the user makes up and down sliding gesture, the system measures the moving distance of the gesture in the vertical direction, maps to the volume increase and decrease, and updates the height of all light columns at the same time; The swipe or tap action in the lyrics area triggers the lyrics display mode switching or annotation jumping, extracts the current lyrics segment and the vocal clarity from the audio in real time, and adjusts the particle aggregation density according to the vocal clarity; The lyrics content is mapped to the particle group, arranged in sequence in the semi-transparent plane or free space, and the particle spacing and cluster size are determined by the particle aggregation density. 6.The intelligent music playing method based on three-dimensional scene interaction of claim 5, wherein: The expression of the rotating angle of the turntable is: wherein: is a rotation angle of the turntable, is a minimum angle of rotation of the turntable, is a maximum angle of rotation of the turntable, is a play progress ratio; The expression of the height of each light column is: Wherein: is the height of each light column, is the device adaptive calibration coefficient, which is set in advance according to the scene size and visual effect; is the loudness level; The expression of the particle aggregation density is: wherein: is a particle concentration density, is a preset least dense particle concentration density, is a preset dense particle concentration density, is a human voice intelligibility.

7. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that: The processor implements the steps of the intelligent music playing method based on three-dimensional scene interaction according to any one of claims 1-6 when executing the computer program.

8. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to implement the steps of the intelligent music playing method based on three-dimensional scene interaction according to any one of claims 1-6.

Citation Information

Patent Citations

  • Interactive music visualization method and device

    CN104732983A

  • Audio playing control method and device, equipment and storage medium

    CN120215870A