Music visualization method and music visualization device

By acquiring the current frame parameter information of the audio signal, determining the target representation mode and generating a visual representation, the limitations of the application scope and real-time performance of existing music visualization methods are solved, achieving a more comprehensive and accurate presentation of music characteristics.

CN121753094APending Publication Date: 2026-03-27HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing music visualization methods are based on MIDI signals or spectrum band processing, which has limited application scope, cannot fully present the characteristics of music, and lacks real-time performance.

Method used

By acquiring the current frame parameter information of the audio signal, the target representation mode is determined, and a visual representation is generated based on the mode, including the analysis of frequency distribution and energy amplitude information, and the representation mode is optimized by combining historical frame information.

Benefits of technology

It achieves comprehensive presentation of music characteristics under high real-time performance, improves the accuracy and reliability of music visualization, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753094A_ABST
    Figure CN121753094A_ABST
Patent Text Reader

Abstract

The invention provides a music visualization method and a music visualization device. The music visualization method comprises the following steps: acquiring a current frame signal of a target audio signal, and determining current frame parameter information based on the current frame signal; determining a target representation mode of the current frame signal based on the current frame parameter information; a visual representation of the current frame signal is generated based on the target representation pattern.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to music processing, and more specifically to a music visualization method and a music visualization device. Background Technology

[0002] With the development of music processing technology, music processing is widely used in the commercial field. The music processing process is also facing higher requirements.

[0003] Currently, music visualization processing is typically based on MIDI (Musical Instrument Digital Interface) signals or on the processing of the music signal's spectral bands. However, MIDI signals are not common audio signal sources, usually requiring music producers to upload or perform complex processing on the music to create a MIDI signal. Therefore, its application is limited and not convenient for users. Spectral band-based visualization processing usually only focuses on some key parts of the music content to ensure real-time performance, which fails to accurately represent all musical characteristics and cannot take into account both high real-time performance and a good overall presentation of musical features.

[0004] Therefore, a music visualization method is needed that not only achieves music visualization but also has good real-time performance, can comprehensively present the characteristics of music at different times, and has high reliability. Summary of the Invention

[0005] According to one aspect of this disclosure, a music visualization method is disclosed, the music visualization method comprising: acquiring a current frame signal of a target audio signal, and determining current frame parameter information based on the current frame signal; determining a target representation mode of the current frame signal based on the current frame parameter information; and generating a visual representation of the current frame signal based on the target representation mode.

[0006] In one embodiment of this disclosure, the current frame parameter information includes at least one of the current frame frequency distribution information and the current frame energy amplitude information.

[0007] In another embodiment of this disclosure, determining the target representation mode of the current frame signal based on the current frame parameter information includes: determining audio feature information corresponding to the current frame signal based on the current frame parameter information; and determining the target representation mode corresponding to the audio feature information among a plurality of preset representation modes based on the audio feature information.

[0008] In another embodiment of this disclosure, the audio feature information includes at least one of overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range.

[0009] In another embodiment of this disclosure, the audio feature information in the low-frequency range includes at least one of low-frequency beat information and low-frequency energy ratio information.

[0010] In another embodiment of this disclosure, the method further includes: acquiring historical frame parameter information of the target audio signal, and wherein determining the target representation mode of the current frame signal based on the current frame parameter information includes: determining audio feature information corresponding to the current frame signal based on the current frame parameter information and historical frame parameter information; and determining the target representation mode corresponding to the audio feature information among a plurality of preset representation modes based on the audio feature information.

[0011] In another embodiment of this disclosure, determining the target representation mode corresponding to the audio feature information among multiple preset representation modes based on the audio feature information includes: comparing the audio feature information with feature information ranges corresponding to the multiple preset representation modes respectively, and forming a comparison result; and determining the preset representation mode corresponding to the audio feature information as the target representation mode based on the comparison result.

[0012] In another embodiment of this disclosure, generating a visual representation of the current frame signal based on a target representation pattern includes: determining visual representation parameters corresponding to the current frame signal based on the target representation pattern and audio feature information; and performing a visual representation of the current frame signal based on the visual representation parameters.

[0013] In another embodiment of this disclosure, the visual representations corresponding to a plurality of preset representation modes are different from each other in at least one of pattern type, pattern color, pattern brightness, and pattern shape.

[0014] According to one aspect of this disclosure, a music visualization device is disclosed, comprising: an audio frame parameter information determination module configured to acquire a current frame signal of a target audio signal and determine current frame parameter information based on the current frame signal; a target representation mode determination module configured to determine a target representation mode of the current frame signal based on the current frame parameter information; and a visual representation generation module configured to generate a visual representation of the current frame signal based on the target representation mode.

[0015] In embodiments of this disclosure, the current frame parameter information includes at least one of the current frame frequency distribution information and the current frame energy amplitude information.

[0016] In another embodiment of this disclosure, the target representation mode determination module includes: an audio feature extraction module configured to determine audio feature information corresponding to the current frame signal based on current frame parameter information; and a target representation mode generation module configured to determine a target representation mode corresponding to the audio feature information from multiple preset representation modes based on the audio feature information.

[0017] In another embodiment of this disclosure, the audio feature information includes at least one of overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range.

[0018] In another embodiment of this disclosure, the audio feature information in the low-frequency range includes at least one of low-frequency beat information and low-frequency energy ratio information.

[0019] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores computer-readable instructions that, when executed by a computer, perform the methods described previously.

[0020] The significance and benefits of this disclosure will become clear from the following description of the embodiments. However, it should be understood that those embodiments are merely examples of how this disclosure may be implemented, and the meaning of the terms used to describe this disclosure is not limited to their specific meaning in describing the embodiments.

[0021] Other systems, methods, features, and advantages of this disclosure will be apparent to those skilled in the art upon review of the following figures and detailed description. All such additional systems, methods, features, and advantages are intended to be included within this specification, within the scope of this disclosure, and protected by the appended claims. Attached Figure Description

[0022] This disclosure can be better understood by referring to the following figures and description. Furthermore, in the figures, the same reference numerals are used throughout different views to refer to corresponding parts.

[0023] Figure 1 A schematic diagram of a music visualization method 100 according to an embodiment of the present disclosure is shown;

[0024] Figure 2 A schematic diagram of process S102 according to an embodiment of the present disclosure is shown;

[0025] Figure 3 A schematic diagram of low-frequency beat information according to an embodiment of the present disclosure is shown;

[0026] Figure 4A schematic diagram of a music visualization method 100 according to another embodiment of the present disclosure is shown;

[0027] Figure 5 An exemplary flowchart of process S103 for generating a visual representation of the current frame signal based on a target representation pattern according to an embodiment of the present disclosure is shown;

[0028] Figure 6 A structural block diagram of a music visualization device 200 according to an embodiment of the present disclosure is shown. Detailed Implementation

[0029] The structure and concepts of this disclosure will be described in more detail below with reference to the accompanying drawings. Details and examples are provided only to facilitate understanding of this disclosure and should not be construed as limiting it in any way.

[0030] Figure 1 A schematic diagram of a music visualization method 100 according to an embodiment of the present disclosure is shown.

[0031] Music visualization refers to the use of computer technology to transform the sound signals of music into visual graphics or animation effects, so as to more intuitively display the rhythm, melody and emotion of music.

[0032] like Figure 1 As shown, according to the music visualization method 100, firstly, in step S101, the current frame signal of the target audio signal is obtained, and the current frame parameter information is determined based on the current frame signal.

[0033] The target audio signal refers to the audio signal of the target music to be visualized. Depending on the actual needs, for example, for the simplicity of streaming, the target audio signal can be a compressed waveform audio format such as AAC and MP3.

[0034] The current frame signal of the target audio signal refers to the target audio signal captured at the current moment. The frame length of the current frame in the audio signal can be approximately 15-30 ms; however, it should be understood that the embodiments of this disclosure are not limited to the frame length of the current frame of the audio signal.

[0035] Current frame parameter information refers to information that reflects the characteristics of the current frame of the audio signal. For example, it can include the characteristics of the current frame of the audio signal in the time-frequency domain.

[0036] For example, parameter information of the current frame can be obtained by performing time-domain analysis and frequency-domain analysis on the current frame; however, the embodiments of this disclosure are not limited thereto.

[0037] For example, the current frame parameter information may include the current frame frequency distribution information and / or the current frame energy amplitude information. However, embodiments of this disclosure are not limited thereto.

[0038] Then, in step S102, the target representation mode of the current frame signal is determined based on the current frame parameter information.

[0039] A target representation pattern refers to the visual representation pattern corresponding to the current frame of an audio signal. It should be understood that a target representation pattern can correspond to a specific representation pattern type, color, brightness, etc.

[0040] For example, the representation mode of the current frame signal can be determined based on the audio feature information corresponding to the current frame signal, which is calculated from the current frame parameter information.

[0041] Then, in step S103, a visual representation of the current frame signal is generated based on the target representation pattern.

[0042] The generation of the visual representation of the current frame signal can be described in detail. For example, a corresponding visual pattern can be presented on the display interface according to the target representation pattern, or it can be achieved through multiple LED light groups located on a speaker, which can flash different colors at corresponding frequencies, or form corresponding patterns / shapes according to the target representation pattern. The embodiments of this disclosure are not limited to the specific performance of the visual representation.

[0043] It should be understood that the music visualization methods and steps in this disclosure can be implemented using various AI (artificial intelligence) methods (e.g., through trained neural networks) or pre-defined integrated algorithms, and can also be implemented using a system combining software and hardware. Embodiments of this disclosure are not limited thereto.

[0044] Based on the above, in this application, the current frame parameter information of each current frame of the target audio signal is determined, and the corresponding target representation mode is determined based on the current frame parameter information for visual display. Therefore, the characteristics of each frame of the target audio signal can be extracted well, and the display mode most suitable for the current frame can be flexibly determined based on these characteristics. This allows the characteristics of each frame of the target audio signal to be presented while achieving good and reliable real-time performance, and also takes into account good music characteristics and high real-time performance.

[0045] In some embodiments, the current frame parameter information includes at least one of the current frame frequency distribution information and the current frame energy amplitude information.

[0046] Current frame frequency distribution information aims to represent the distribution characteristics of the current frame of an audio signal at different frequencies. For example, the current frame can be processed in the frequency domain to obtain its spectrum in multiple different frequencies / frequency ranges (such as the low-frequency range of 50 Hz to 200 Hz; the mid-frequency range of 200 Hz to 4000 Hz; and the high-frequency range of ≥4000 Hz), which can be used as the current frame frequency distribution information.

[0047] However, it should be understood that the current frame frequency distribution information is not limited to the above forms.

[0048] Current frame energy amplitude information refers to the amplitude value of the current frame of the audio signal. For example, the current frame of the audio signal can be processed in the time domain to obtain the amplitude value in the current frame.

[0049] Based on the above, in this application, by setting the current frame parameter information to include at least one of the current frame frequency distribution information and the current frame energy amplitude information, the current frame parameter information can well reflect the characteristics of the current frame of the audio signal in the time domain and / or frequency domain. This is beneficial for well determining its target representation mode and accurately presenting the characteristics of the current frame.

[0050] In some embodiments, the process S102 of determining the target representation mode of the current frame signal based on the current frame parameter information can be described in more detail. Figure 2 A schematic diagram of process S102 according to an embodiment of the present disclosure is shown;

[0051] refer to Figure 2 The above process S102 includes: First, in step S1021, determining the audio feature information corresponding to the current frame signal based on the current frame parameter information.

[0052] Audio feature information refers to information representing the musical characteristics of the current frame of an audio signal. For example, it represents the tempo characteristics, volume characteristics, frequency response characteristics, and frequency distribution characteristics of the current frame. It should be noted that embodiments of this disclosure are not limited to a specific composition of audio feature information.

[0053] For example, the current frame parameter information can be directly used as audio feature information, or the current frame parameter information can be further processed and analyzed to obtain audio feature information.

[0054] Then, in step S1022, based on the audio feature information, a target representation mode corresponding to the audio feature information is determined from multiple preset representation modes.

[0055] For example, preset display modes can be imported by the system or set manually by the user. Preset display modes can include multiple different display modes, such as silent mode, pure music mode, strong beat mode, complex music mode, climax mode, and transition mode.

[0056] For example, the above pattern can be described in detail as follows.

[0057] The silent mode (Mode A) indicates that there is very little or no sound in the music. Silent audio signals in the intro and outro usually fall into this category.

[0058] Pure music mode (Mode B) indicates that the music is mainly composed of the singer's voice without many other instruments or strong beats.

[0059] The strong beat mode (mode C) indicates a situation where a strong beat (bass or drum beat) is detected on a simple melody.

[0060] Complex music patterns (Pattern D) represent situations where the music lacks strong bass but conveys complex information. For example, when a choir sings with the lead singer.

[0061] Climax mode (E mode) indicates a situation where both strong beats and very complex musical information are presented in the music; it usually indicates the most powerful part of the music.

[0062] A transitional mode (Mode F) signifies that the music is currently undergoing a change in its properties. The mode may revert to its beginning, switch to another mode, or reach a new climax. Transitional modes are only temporary.

[0063] It should be understood that the above description provides only some examples of representation patterns. The embodiments of this disclosure are not limited to a specific composition of a preset representation pattern.

[0064] For example, the process of determining the target representation mode can be achieved by comparing audio feature information with feature information ranges corresponding to multiple preset representation modes; and determining the preset representation mode corresponding to the audio feature information as the target representation mode based on the comparison results. However, the process of determining the target representation mode is not limited to this.

[0065] For example, if the overall frequency distribution information is in a uniform distribution state (for example, the uniform distribution state indicates that the proportion of low-frequency signals and the proportion of high-frequency signals in the overall frequency distribution information are both less than a first proportion threshold), and the audio feature information (low-frequency beat information and low-frequency energy ratio information) in the low-frequency range is less than a low-frequency preset first threshold, and the audio feature information in the high-frequency range is less than a high-frequency preset first threshold, then it is determined to be a silent mode (mode A).

[0066] For example, if the overall frequency distribution information is uniformly distributed, and the audio feature information in the low-frequency range is greater than the low-frequency preset first threshold, and the audio feature information in the high-frequency range is greater than the high-frequency preset first threshold, then it is determined to be a pure music mode (mode B).

[0067] For example, if the overall frequency distribution information is in a low-frequency dominant distribution state (for example, the low-frequency dominant distribution state indicates that the proportion of low-frequency signals in the overall frequency distribution information is greater than a second proportion threshold, wherein the second proportion threshold is greater than the aforementioned first proportion threshold), and the audio feature information in the low-frequency range is greater than a low-frequency preset second threshold (the low-frequency preset second threshold is greater than a low-frequency preset first threshold), and the audio feature information in the high-frequency range is less than a high-frequency preset second threshold (the high-frequency preset second threshold is greater than a high-frequency preset first threshold), then it is determined to be a strong beat mode (mode C).

[0068] For example, if the overall frequency distribution information is in a high-frequency dominant distribution state (e.g., the high-frequency dominant distribution state indicates that the proportion of high-frequency signals in the overall frequency distribution information is greater than a second proportion threshold, wherein the second proportion threshold is greater than the aforementioned first proportion threshold), and the audio feature information in the low-frequency range is less than a low-frequency preset second threshold, and the audio feature information in the high-frequency range is greater than a high-frequency preset second threshold, then it is determined to be a complex mode (mode D).

[0069] For example, if the overall frequency distribution information is in a special uniform distribution state (for example, the special uniform distribution state indicates that the proportion of high-frequency signals in the overall frequency distribution information and the proportion of low-frequency signals in the overall frequency distribution information are both greater than a second proportion threshold, wherein the second proportion threshold is greater than the first proportion threshold mentioned above), and the audio feature information in the low-frequency range is greater than the low-frequency preset second threshold, and the audio feature information in the high-frequency range is also greater than the high-frequency preset second threshold, then it is determined to be a climax mode (mode E).

[0070] It should be noted that the above is only an exemplary process for determining the target representation pattern, and the embodiments of this disclosure are not limited thereto.

[0071] Based on the above, by determining the audio feature information corresponding to the current frame signal based on the current frame parameter information, and further determining the target representation mode corresponding to the audio feature information from multiple preset representation modes based on the audio feature information, this application can flexibly select an appropriate representation mode according to the audio characteristics of the current frame, thereby comprehensively and holistically realizing the visualization process of the audio characteristics of the current frame, achieving more accurate and comprehensive visual display, and improving user experience.

[0072] In some embodiments, the audio feature information includes at least one of overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range.

[0073] Overall frequency distribution information aims to characterize the frequency characteristics of the current frame of an audio signal. For example, the frequency distribution information of the current frame can be used directly as overall frequency distribution information, or the frequency distribution information of the current frame can be further processed to generate overall frequency distribution information.

[0074] Audio feature information in the low-frequency range aims to characterize the properties of the current frame of an audio signal in the low-frequency range. For example, audio feature information in the low-frequency range may include low-frequency beat information and low-frequency energy ratio information.

[0075] Specifically, for example, 50Hz to 200Hz is generally considered low frequency, typically encompassing strong drum beats and other bass instruments. Drum beats are characterized by this frequency band, thus allowing the extraction of beat characteristics with relatively large gaps from those frequencies.

[0076] Audio feature information in the high-frequency range is intended to characterize the properties of the current frame of an audio signal in the high-frequency range.

[0077] Specifically, frequencies of 4000Hz and higher are considered high frequencies, encompassing some of the fundamental overtones of sound, which are then considered the texture or timbre of the sound. Music visualization methods can focus on even higher frequencies because they contain percussion sounds produced by certain instruments. For example, this method can calculate prominent high-frequency sounds produced by percussion instruments, such as the metallic sound from the metal strings of a snare drum.

[0078] In some embodiments, the audio feature information includes all of the overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range, in order to better represent the characteristics of the current frame of the audio signal.

[0079] In some embodiments, the audio feature information further includes audio feature information in the mid-frequency range, which is intended to characterize the characteristics of the current frame of the audio signal in the mid-frequency range.

[0080] For example, 200Hz–4000Hz is considered the midrange, encompassing the fundamental frequencies and major overtones of vocals and other instruments. These frequencies will be processed into smooth, continuous energy over time. Filters will be used to suppress unexpected attacks. By utilizing filters that suppress abrupt changes in energy, the system can preserve vocal detail while reducing the presence of percussion instruments.

[0081] It should be understood that the audio feature information may also include other content as needed, and the embodiments of this disclosure are not limited thereto.

[0082] Based on the above, by setting the audio feature information to include at least one of the overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range, the audio feature information can reflect the characteristics of the current frame of the audio signal in one or more aspects of the overall frequency distribution characteristics, low-frequency characteristics, and high-frequency characteristics, thus promoting the subsequent corresponding visual display based on the audio feature information.

[0083] In some embodiments, the audio feature information in the low-frequency range includes at least one of low-frequency beat information and low-frequency energy ratio information.

[0084] Low-frequency beat information refers to the beat characteristics of the current frame of an audio signal in the low-frequency range.

[0085] Specifically, for example, energy amplitude information corresponding to the low-frequency range of the current frame of the audio signal can be obtained, and low-frequency beat information can be determined based on the energy amplitude information in the low-frequency range. For example, local maxima of energy amplitude information in the low-frequency range can be found, and single-point beat positions can be assigned. Then, for example, output shaping can be performed to make the energy amplitude of the beat have a uniform value or within a preset amplitude range, and the generated low-frequency beat information (such as a beat energy amplitude-time diagram) can be generated.

[0086] Figure 3 A schematic diagram of low-frequency beat information according to an embodiment of the present disclosure is shown. The low-frequency beat information is presented in the form of a beat energy amplitude-time diagram. Specifically, Figure 3 The x-axis in the diagram represents time, and Figure 3 The y-axis represents the energy amplitude. The dashed line shows the current frame energy amplitude information in the low-frequency range graph F1, and the solid line shows the beat energy amplitude-time graph F2. The process algorithm finds local maxima, assigns single-point beat positions, and reshapes the beat output into a uniform decay pattern. It also sets a minimum gap between beats to avoid unwanted continuous triggering of bass beat parameters due to data noise.

[0087] Low-frequency energy ratio information refers to the ratio of audio energy in the low-frequency range to the total audio energy. This low-frequency energy ratio information can be determined based on the current frame frequency distribution information.

[0088] Based on the above, by setting the audio feature information in the low-frequency range to include at least one of low-frequency beat information and low-frequency energy ratio information, the audio feature information obtained in the low-frequency range can well reflect the characteristics of the current frame in low-frequency beat and low-frequency ratio, which is beneficial for determining the corresponding representation mode based on these characteristics.

[0089] Figure 4A schematic diagram of a music visualization method 100 according to another embodiment of the present disclosure is shown.

[0090] refer to Figure 4 In some embodiments, the music visualization method further includes step S101' of obtaining historical frame parameter information of the target audio signal.

[0091] The historical frame parameter information of the target audio signal represents the parameter information of the audio signal in multiple preset frames preceding the current frame, aiming to reflect the audio characteristics of audio frames adjacent to the current frame. For example, it can include the corresponding parameter information of multiple historical frames, or it can be comprehensive parameter information obtained by processing the parameter information of multiple historical frames.

[0092] It should be understood that the operations of steps S101 and S101' can be performed in parallel or sequentially, and this document makes no restrictions on this.

[0093] The process S102, which determines the target representation mode of the current frame signal based on the current frame parameter information, may include: First, in step S1021', determining the audio feature information corresponding to the current frame signal based on the current frame parameter information and historical frame parameter information. Then, in step S1022, determining the target representation mode corresponding to the audio feature information from multiple preset representation modes based on the audio feature information.

[0094] The process of determining audio feature information corresponding to the current frame signal based on current frame parameter information and historical frame parameter information can be described in detail. For example, the current frame parameter information may include current frame energy amplitude information, and the historical frame parameter information may include historical frame energy amplitude information representing the average energy amplitude of historical frames. Additionally, the current frame parameter information may include audio feature information in the high-frequency range (such as high-frequency energy change rate information), wherein the high-frequency energy change rate information can be calculated using the current frame energy amplitude information and the historical frame energy amplitude information. However, it should be understood that embodiments of this disclosure are not limited thereto.

[0095] The process of step S1022 has been described in detail above and will not be repeated here.

[0096] Based on the above, in this application, by obtaining the historical frame parameter information of the target audio signal, and determining the audio feature information corresponding to the current frame based on the current frame parameter information and the historical frame parameter information, the determined audio feature signal can simultaneously consider the characteristics of the current frame and the characteristics of the historical frames adjacent to the current frame, so that the representation form suitable for the current frame can be determined more accurately, and the visual representation can take into account the consistency with the historical frames and the characteristics of the current frame.

[0097] In some embodiments, the above-described process S1022, which determines the target representation pattern corresponding to the audio feature information among a plurality of preset representation patterns based on the audio feature information, can be described more specifically. Figure 4 An exemplary flowchart of a process 1022 for determining a target representation pattern corresponding to audio feature information among a plurality of preset representation patterns according to an embodiment of the present disclosure is shown.

[0098] like Figure 4 As shown, firstly, in step S1022-1, the audio feature information is compared with the feature information range corresponding to multiple preset representation modes, and a comparison result is formed.

[0099] The feature information range corresponding to the preset representation mode refers to the preset information value range of the relevant feature information corresponding to the preset representation mode.

[0100] For example, the process of comparing audio feature information with feature information ranges corresponding to multiple preset representation patterns can be described in detail. For example, the audio feature information may include multiple sub-feature information. For example, for each preset representation pattern, the sub-feature information can be compared with a preset sub-feature information range corresponding to that preset representation pattern, and the number of cases where the sub-feature information falls within the preset sub-feature information range can be counted. For example, the number of sub-feature information falling within the corresponding sub-feature information range can be counted, and the count value can be used as the comparison result.

[0101] However, it should be understood that the above is merely an example of comparison and comparison results, and the embodiments of this disclosure are not limited thereto.

[0102] Then, in step S1022-2, the preset representation mode corresponding to the audio feature information is determined as the target representation mode based on the comparison result.

[0103] For example, if the comparison result indicates the number of sub-features falling within the range of the corresponding sub-feature information, then based on the comparison result corresponding to each preset representation mode, the preset representation mode with the maximum value in the comparison result (i.e., the preset representation mode that is most consistent with the audio feature information of the current frame) is determined as the target representation mode.

[0104] Based on the above, by comparing the audio feature information with the feature information range corresponding to the preset representation mode, and based on the comparison result, the preset representation mode corresponding to the audio feature information is determined as the target representation mode, so that the determined target representation mode can be well applied to the audio feature information, thereby accurately and reliably displaying the characteristics of the current frame of the audio signal, and improving the accuracy and reliability of music visualization.

[0105] In some embodiments, the process S103 of generating a visual representation of the current frame signal based on the target representation pattern can be described more specifically.

[0106] Figure 5 An exemplary flowchart of process S103 for generating a visual representation of the current frame signal based on a target representation pattern according to an embodiment of the present disclosure is shown.

[0107] like Figure 5 As shown, firstly, in step S1031, visual representation parameters corresponding to the current frame signal are determined based on the target representation mode and audio feature information.

[0108] Visual representation parameters refer to the parameters used to generate music visualizations. Depending on the actual performance, representation parameters can be screen display parameters (such as the shape, pattern, and color to be displayed), or they can be lighting control information that controls the illumination of the LED array (such as lighting color, lighting pattern, and lighting frequency).

[0109] The process of determining the visual representation parameters corresponding to the current frame signal based on the target representation pattern and audio feature information may include, for example, determining the representation form (such as a pattern, color, etc.) corresponding to the target representation pattern. For example, in quiet mode, its musical characteristics are represented by colored solid lines; in strong beat mode, for example, beat points are represented by impact waveforms, or beat points are represented by circular patterns flashing at a fixed frequency. Then, based on the representation form and audio feature information, visual representation parameters are generated, for example, determining beat points based on impact waveforms in the strong beat mode, and determining the specific display frequency and amplitude of the beat points based on specific audio feature information (low-frequency beat information).

[0110] Next, in step S1032, the visual representation of the current frame signal is performed based on the visual representation parameters.

[0111] It should be understood that the embodiments of this disclosure are not limited to any particular display device.

[0112] Based on the above, in this application, visual representation parameters corresponding to the current frame signal are determined according to the target representation mode and audio feature information, and the arrangement of the visual representation of the current frame signal is performed based on the visual representation parameters, so that the target representation mode and audio feature information corresponding to the current frame signal are considered simultaneously during display, thereby enabling a comprehensive and reliable display.

[0113] In some embodiments, the visual representations corresponding to a plurality of preset representation modes are different from each other in at least one of pattern type, pattern color, pattern brightness, and pattern shape.

[0114] Pattern type is designed to characterize the type of a visual pattern, and pattern type can be, for example, a single point, an array of points, a single line, multiple lines, or one or more closed figures.

[0115] Pattern color is designed to represent the colors that a visual pattern possesses; it can be a single color or a combination of one or more preset colors.

[0116] Pattern brightness is designed to characterize the brightness of a visual pattern; it can be set by the system or preset by the user.

[0117] Pattern shape is designed to characterize the shape of a visual pattern. For example, if the pattern type is a single line, the pattern shape can be a solid line, a dashed line, or a dotted line. Additionally, if the pattern type is a closed figure, the pattern shape can be a rectangle, a square, a circle, an ellipse, etc.

[0118] Based on the above, by arranging the visual representations corresponding to multiple preset representation modes differently in at least one of pattern type, pattern color, pattern brightness, and pattern shape, the presentation of music visualization in this application is more diversified, and each different representation mode can choose the presentation form that is most conducive to displaying its audio characteristics, thus improving the richness of music visualization and enhancing the user experience.

[0119] According to another aspect of this disclosure, a music visualization device is provided. Figure 6 A structural block diagram of a music visualization device 200 according to an embodiment of the present disclosure is shown.

[0120] like Figure 6 As shown, the music visualization device 200 provided by this embodiment includes: an audio frame parameter information determination module 210, a target representation mode determination module 220, and a visual representation generation module 230. The music visualization device 200 can perform... Figure 1 The method shown.

[0121] The audio frame parameter information determination module 210 is configured to perform Figure 1 The operation in step S101 can acquire the current frame signal of the target audio signal and determine the current frame parameter information based on the current frame signal.

[0122] The target audio signal refers to the audio signal of the target music to be visualized. Depending on the actual needs, for example, for the simplicity of streaming, the target audio signal can be a compressed waveform audio format such as AAC and MP3.

[0123] The current frame signal of the target audio signal refers to the target audio signal captured at the current moment. The frame length of the current frame in the audio signal can be approximately 15-30 ms; however, it should be understood that the embodiments of this disclosure are not limited to the frame length of the current frame of the audio signal.

[0124] Current frame parameter information refers to information that reflects the characteristics of the current frame of the audio signal. For example, it can include the characteristics of the current frame of the audio signal in the time-frequency domain.

[0125] For example, parameter information of the current frame can be obtained by performing time-domain analysis and frequency-domain analysis on the current frame; however, the embodiments of this disclosure are not limited thereto.

[0126] For example, the current frame parameter information may include the current frame frequency distribution information and / or the current frame energy amplitude information. However, embodiments of this disclosure are not limited thereto.

[0127] Target representation pattern determination module 220 is configured to execute Figure 1 The operation in step S102 can determine the target representation mode of the current frame signal based on the current frame parameter information.

[0128] A target representation pattern refers to the visual representation pattern corresponding to the current frame of an audio signal. It should be understood that a target representation pattern can correspond to a specific representation pattern type, color, brightness, etc.

[0129] For example, the representation mode of the current frame signal can be determined based on the audio feature information corresponding to the current frame signal, which is calculated from the current frame parameter information.

[0130] The visual representation generation module 230 is configured to perform Figure 1 The operation in step S103 can generate a visual representation of the current frame signal based on the target representation mode.

[0131] The generation of the visual representation of the current frame signal can be described in detail. For example, a corresponding visual pattern can be presented on the display interface according to the target representation pattern, or it can be achieved through multiple LED light groups located on a speaker, which can flash different colors at corresponding frequencies, or form corresponding patterns / shapes according to the target representation pattern. The embodiments of this disclosure are not limited to the specific performance of the visual representation.

[0132] Based on the above, in this application, the current frame parameter information of each current frame of the target audio signal is determined, and the corresponding target representation mode is determined based on the current frame parameter information for visual display. Therefore, the characteristics of each frame of the target audio signal can be extracted well, and the display mode most suitable for the current frame can be flexibly determined based on these characteristics. This allows the characteristics of each frame of the target audio signal to be presented while achieving good and reliable real-time performance, and also takes into account good music characteristics and high real-time performance.

[0133] In some embodiments, the current frame parameter information includes at least one of the current frame frequency distribution information and the current frame energy amplitude information.

[0134] Current frame frequency distribution information aims to represent the distribution characteristics of the current frame of an audio signal at different frequencies. For example, the current frame can be processed in the frequency domain to obtain its spectrum in multiple different frequencies / frequency ranges, which can be used as the current frame frequency distribution information. However, it should be understood that the current frame frequency distribution information is not limited to the above form.

[0135] Current frame energy amplitude information refers to the amplitude value of the current frame of the audio signal. For example, the current frame of the audio signal can be processed in the time domain to obtain the amplitude value in the current frame.

[0136] Based on the above, in this application, by setting the current frame parameter information to include at least one of the current frame frequency distribution information and the current frame energy amplitude information, the current frame parameter information can well reflect the characteristics of the current frame of the audio signal in the time domain and / or frequency domain. This is beneficial for well determining its target representation mode and accurately presenting the characteristics of the current frame.

[0137] In some embodiments, the target representation pattern determination module 220 includes an audio feature extraction module 221 and a target representation pattern generation module 222.

[0138] The audio feature extraction module 221 is configured to perform Figure 2 The operation in step S1021 can determine the audio feature information corresponding to the current frame signal based on the current frame parameter information.

[0139] Audio feature information refers to information representing the musical characteristics of the current frame of an audio signal. For example, it represents the tempo characteristics, volume characteristics, frequency response characteristics, and frequency distribution characteristics of the current frame. It should be noted that embodiments of this disclosure are not limited to a specific composition of audio feature information.

[0140] For example, the current frame parameter information can be directly used as audio feature information, or the current frame parameter information can be further processed and analyzed to obtain audio feature information.

[0141] The target representation pattern generation module 222 is configured to execute Figure 2 The operation of step S1022 can determine the target representation mode corresponding to the audio feature information among multiple preset representation modes based on the audio feature information.

[0142] For example, preset display modes can be imported by the system or manually set by the user. Preset display modes may include multiple different modes, such as silent mode, pure music mode, strong beat mode, complex music mode, climax mode, and transition mode. The embodiments of this disclosure are not limited to a specific composition of preset display modes.

[0143] For example, the process of determining the target representation mode can be achieved by comparing audio feature information with feature information ranges corresponding to multiple preset representation modes; and determining the preset representation mode corresponding to the audio feature information as the target representation mode based on the comparison results. However, the process of determining the target representation mode is not limited to this.

[0144] Based on the above, by determining the audio feature information corresponding to the current frame signal based on the current frame parameter information, and further determining the target representation mode corresponding to the audio feature information from multiple preset representation modes based on the audio feature information, this application can flexibly select an appropriate representation mode according to the audio characteristics of the current frame, thereby comprehensively and holistically realizing the visualization process of the audio characteristics of the current frame, achieving more accurate and comprehensive visual display, and improving user experience.

[0145] In some embodiments, the audio feature information includes at least one of overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range.

[0146] Overall frequency distribution information aims to characterize the frequency characteristics of the current frame of an audio signal. For example, the frequency distribution information of the current frame can be used directly as overall frequency distribution information, or the frequency distribution information of the current frame can be further processed to generate overall frequency distribution information.

[0147] Audio feature information in the low-frequency range aims to characterize the properties of the current frame of an audio signal in the low-frequency range. For example, audio feature information in the low-frequency range may include low-frequency beat information and low-frequency energy ratio information.

[0148] Audio feature information in the high-frequency range is intended to characterize the properties of the current frame of an audio signal in the high-frequency range.

[0149] In some embodiments, the audio feature information includes all of the overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range, in order to better represent the characteristics of the current frame of the audio signal.

[0150] It should be understood that the audio feature information may also include other content as needed, and the embodiments of this disclosure are not limited thereto.

[0151] Based on the above, by setting the audio feature information to include at least one of the overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range, the audio feature information can reflect the characteristics of the current frame of the audio signal in one or more aspects of the overall frequency distribution characteristics, low-frequency characteristics, and high-frequency characteristics, thus promoting the subsequent corresponding visual display based on the audio feature information.

[0152] In some embodiments, the audio feature information in the low-frequency range includes at least one of low-frequency beat information and low-frequency energy ratio information.

[0153] Low-frequency beat information refers to the beat characteristics of the current frame of an audio signal in the low-frequency range.

[0154] Low-frequency energy ratio information refers to the ratio of audio energy in the low-frequency range to the total audio energy. This low-frequency energy ratio information can be determined based on the current frame frequency distribution information.

[0155] Based on the above, by setting the audio feature information in the low-frequency range to include at least one of low-frequency beat information and low-frequency energy ratio information, the audio feature information obtained in the low-frequency range can well reflect the characteristics of the current frame in low-frequency beat and low-frequency ratio, which is beneficial for determining the corresponding representation mode based on these characteristics.

[0156] In some embodiments, the music visualization device may perform the methods described above and achieve the functions described above.

[0157] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is also provided, on which computer-readable instructions are stored, which, when executed by a computer, perform the method described above.

[0158] The program portion of a technology can be considered a "product" or "artifact" existing in the form of executable code and / or related data, which is embodied or implemented by a computer-readable medium. Tangible, permanent storage media may include internal memory or storage devices used by any computer, processor, or similar device or associated module. For example, various semiconductor memories, magnetic tape drives, disk drives, or any similar devices capable of providing storage functionality for software.

[0159] All software, or parts thereof, can sometimes communicate via networks such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. For example, loading software from a server or host computer to a hardware environment of a computer environment, or another computer environment that implements the system, or a system with similar functionality associated with providing the information required for image processing. Therefore, another medium capable of transmitting software elements can also serve as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., to be propagated through cables, optical fibers, or air. Physical media used to carry waves (such as cables, wireless connections, or fiber optic cables) can also be considered as media used to carry software. In this context, unless a tangible “storage” medium is defined, other terms referring to a “readable medium” for a computer or machine mean the medium involved in the processor executing any instructions.

[0160] While various embodiments of this disclosure have been described, it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of this disclosure. Therefore, this disclosure is not limited except as provided in the appended claims and their equivalents.

Claims

1. A music visualization method, comprising: Obtain the current frame signal of the target audio signal, and determine the current frame parameter information based on the current frame signal; The target representation mode of the current frame signal is determined based on the current frame parameter information; The visual representation of the current frame signal is generated based on the target representation pattern.

2. The music visualization method according to claim 1, wherein the current frame parameter information includes at least one of the current frame frequency distribution information and the current frame energy amplitude information.

3. The music visualization method according to claim 1, wherein determining the target representation mode of the current frame signal based on the current frame parameter information includes: Based on the current frame parameter information, determine the audio feature information corresponding to the current frame signal; Based on the audio feature information, a target representation pattern corresponding to the audio feature information is determined from multiple preset representation patterns.

4. The music visualization method according to claim 3, wherein the audio feature information includes at least one of overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range.

5. The music visualization method according to claim 4, wherein the audio feature information in the low-frequency range includes at least one of low-frequency beat information and low-frequency energy ratio information.

6. The music visualization method according to claim 1, further comprising: Obtain historical frame parameter information of the target audio signal, and The determination of the target representation mode of the current frame signal based on the current frame parameter information includes: The audio feature information corresponding to the current frame signal is determined based on the current frame parameter information and the historical frame parameter information. Based on the audio feature information, a target representation pattern corresponding to the audio feature information is determined from multiple preset representation patterns.

7. The music visualization method according to claim 3, wherein determining the target representation pattern corresponding to the audio feature information among the plurality of preset representation patterns based on the audio feature information comprises: The audio feature information is compared with the feature information ranges corresponding to the plurality of preset representation modes, and a comparison result is formed; Based on the comparison results, the preset representation pattern corresponding to the audio feature information is determined as the target representation pattern.

8. The music visualization method according to claim 3, wherein generating a visual representation of the current frame signal based on the target representation pattern comprises: The visual representation parameters corresponding to the current frame signal are determined based on the target representation pattern and the audio feature information; The visual representation of the current frame signal is performed based on the visual representation parameters.

9. The music visualization method according to claim 1, wherein the visual representations corresponding to the plurality of preset representation modes are different from each other in at least one of pattern type, pattern color, pattern brightness, and pattern shape.

10. A music visualization device, comprising: An audio frame parameter information determination module is configured to acquire the current frame signal of the target audio signal and determine the current frame parameter information based on the current frame signal. A target representation mode determination module is configured to determine the target representation mode of the current frame signal based on the current frame parameter information. A visual representation generation module is configured to generate a visual representation of the current frame signal based on the target representation pattern.

11. The music visualization device according to claim 10, wherein the current frame parameter information includes at least one of current frame frequency distribution information and current frame energy amplitude information.

12. The music visualization apparatus according to claim 10, wherein the target representation mode determination module comprises: An audio feature extraction module is configured to determine audio feature information corresponding to the current frame signal based on the current frame parameter information. A target representation pattern generation module is configured to determine a target representation pattern corresponding to the audio feature information from multiple preset representation patterns based on the audio feature information.

13. The music visualization device according to claim 12, wherein the audio feature information includes at least one of overall frequency distribution information, audio feature information in the low-frequency range, and audio feature information in the high-frequency range.

14. The music visualization device according to claim 13, wherein the audio feature information in the low-frequency range includes at least one of low-frequency beat information and low-frequency energy ratio information.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, and when the instructions are executed by a computer, performs the method according to any one of claims 1 to 9.