AI voice interaction CD play control method and device

By adopting AI voice interaction technology in CD players, the problem that traditional CD players cannot understand user voice commands is solved, and the satisfaction of user personalized audio experience is achieved.

CN120032641APending Publication Date: 2025-05-23SHENZHEN ZHONGLIN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510111055.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional CD players lack intelligence, cannot understand users' voice commands, and cannot meet users' needs for personalized audio experience.

Method used

Using AI voice interaction technology, users' voice signals are collected through microphones, digital processing and preprocessing, voice signal characteristic parameters are extracted, and voice recognition model is input into the speech recognition model constructed by deep learning algorithms, classified recognition is performed, and text control instructions are generated.

Benefits of technology

It realizes an accurate understanding of user voice commands, automatically playback control and music recommendations based on user needs, meeting users' needs for convenient, intelligent and personalized audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032641A_ABST
    Figure CN120032641A_ABST
Patent Text Reader

Abstract

The invention discloses an AI voice interaction CD playing control method and equipment, and the method comprises the steps: collecting a voice signal of a user through calling a microphone of an AI voice recognition module, and carrying out the digital processing of the collected voice signal into a digital audio signal; preprocessing the digital audio signal; voice signal feature parameter extraction is carried out on the preprocessed digital audio signals, voice signal feature parameters comprise Mel-frequency cepstral coefficient feature vectors, and acoustic features of the voice signals are represented through the voice signal feature parameters; and inputting the extracted voice signal feature parameters into a pre-trained voice recognition model constructed based on a deep learning algorithm, and classifying and recognizing the input voice signal feature parameters through the voice recognition model to obtain a text control instruction for the CD playing equipment. According to the method, the voice instruction of the user can be accurately understood, playing control and music recommendation are automatically carried out according to personalized requirements and preferences of the user, and the pursuit of personalized audio experience of the user is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an AI voice interactive CD playback control method and device, belonging to the technical field of players. Background Art

[0002] A CD player, also known as a laser turntable or laser player, is an intelligent high-fidelity stereo sound device controlled by a microcomputer. It uses advanced laser technology, digital technology, computer technology and various new components. It has the advantages of high-density recording, long playback time, easy operation and fast song selection. It can realistically reproduce the recorded content with clear layers and a sense of presence.

[0003] The operation of traditional CD players mainly relies on the physical buttons on the body, and users need to manually operate to achieve functions such as play, pause, song selection, and volume adjustment. This operation method is inconvenient in many scenarios. For example, when the user's hands are busy with other matters or the player is far away, it is difficult to perform accurate operations. In addition, traditional CD players lack intelligence and cannot understand the user's voice commands. They cannot automatically control playback and recommend music according to the user's personalized needs and preferences, and it is difficult to meet the modern user's pursuit of convenient, intelligent, and personalized audio experience. Summary of the invention

[0004] To this end, the present invention provides an AI voice interactive CD playback control method and device to solve the problem that traditional CD players lack intelligence, cannot understand user voice commands, and cannot meet users' personalized audio experience.

[0005] In order to achieve the above object, the present invention provides the following technical solution: an AI voice interactive CD playback control method, comprising the following steps:

[0006] The microphone of the AI ​​speech recognition module is called to collect the user's voice signal, and the collected voice signal is digitized and processed into a digital audio signal;

[0007] Preprocessing the digital audio signal, wherein the preprocessing includes noise reduction, filtering, and frame division;

[0008] Extracting speech signal feature parameters from the preprocessed digital audio signal, wherein the speech signal feature parameters include Mel-frequency cepstral coefficient feature vectors, and representing the acoustic features of the speech signal through the speech signal feature parameters;

[0009] The extracted voice signal feature parameters are input into a pre-trained voice recognition model constructed based on a deep learning algorithm, and the input voice signal feature parameters are classified and recognized by the voice recognition model to obtain text control instructions for the CD player.

[0010] As a preferred solution of the AI ​​voice interactive CD player control method, during the preprocessing of the digital audio signal:

[0011] Define an array to store the sample values ​​of the digital audio signal;

[0012] The array format is: int audioSamples[BUFFER_SIZE], where BUFFER_SIZE represents the array size of the cached audio sample values;

[0013] Taking the sampling value array and the array size of the digital audio signal as parameters, by looping through the array, taking the set window size as a unit, calculating the average value of the sampling values ​​in each window;

[0014] A nested loop is constructed, wherein the nested loop includes an outer loop and an inner loop; the sliding position of the window is controlled by the outer loop, and the sampling values ​​in the window are accumulated and averaged by the inner loop; and the sampling value at the center of the window is replaced by the average value.

[0015] As a preferred solution of the AI ​​voice interactive CD player control method, the outer loop starts from half of the window size and ends at the end of the array minus half of the window size, so that each sampling point that needs filtering is located at the center of the window;

[0016] The window range of the inner loop is centered on the current outer loop position i, and is extended by windowSize / 2 sampling points on the left and right. WindowSize is the window size of the mean filter.

[0017] As a preferred solution for the AI ​​voice interactive CD player control method, in the voice recognition model constructed based on the deep learning algorithm, the output values ​​of the voice input signal are multiplied by the corresponding weights and then added together. The obtained results are added with the bias value corresponding to the voice output signal, and finally the final output result of the voice output signal is obtained through the activation function.

[0018] As a preferred solution of the AI ​​voice interactive CD player control method, the expression of the voice recognition model is:

[0019]

[0020] Where y(n) is the speech signal output by the model for the nth class, and w i is the weight of speech output signal i, x i is the speech input signal i, b i is the bias value corresponding to the speech output signal i.

[0021] The present invention also provides an AI voice interactive CD playback control device, which adopts the above-mentioned AI voice interactive CD playback control method, including a main control module, an AI voice recognition module, a CD player audio link module, an audio playback module and a storage module;

[0022] The AI ​​voice recognition module includes a voice recognition chip and a microphone, the microphone and the voice recognition chip are electrically connected, the voice recognition chip and the main control module are electrically connected, and the AI ​​voice recognition module is used to collect the user's voice control instructions;

[0023] The CD player audio link module includes a movement, a drive chip and an audio decoding chip; a connection is established between the movement and the drive chip, a connection is established between the drive chip and the audio decoding chip, and a connection is established between the audio decoding chip and the main control module; the movement is used to scan the CD disc information through a laser head; the drive chip is used to read audio data; the audio decoding chip is used to decode the read audio data into an audio signal;

[0024] A connection is established between the audio playback module and the main control module. The audio playback module is provided with an audio power amplifier chip. The audio playback module is used to play the original audio signal through the audio power amplifier chip and the configured speaker;

[0025] A connection is established between the storage module and the main control module, and the storage module is used to randomly store the audio signal converted by decoding the audio data.

[0026] As a preferred solution for the AI ​​voice interactive CD playback control device, it also includes a dual-mode wireless module, a connection is established between the dual-mode wireless module and the main control module, and the dual-mode wireless module is used for wireless data transmission and control of the CD playback device.

[0027] As a preferred solution of the AI ​​voice interactive CD playback control device, it also includes a mobile server module, the mobile server module establishes a connection with the main control module through the dual-mode wireless module, and the mobile server module is used for the user to send playback control instructions to the CD playback device through the mobile terminal;

[0028] The mobile server module uses a trained speech recognition model built based on a deep learning algorithm for comparison and matching to obtain a recognition result, and transmits the recognition result back to the main control module. After the main control module parses the recognition result, it performs operations or displays according to the parsed result to complete the interaction.

[0029] As a preferred solution for the AI ​​voice interactive CD playback control device, it also includes a display module, which is configured with a TFT touch color screen. The display module is connected to the main control module, and the display module is used to display the music lyrics data corresponding to the played audio signal.

[0030] As a preferred solution for the AI ​​voice interactive CD playback control device, it also includes a power module, which is configured with a battery, a 3.3V voltage regulator chip, a charging management chip and a charging interface;

[0031] The battery is used to power the entire CD player; the charging interface is used to connect an external power adapter to charge the battery; the charging management chip is used to regulate the charging process; the 3.3V voltage regulator chip is used to convert the voltage input from the battery or the charging interface into a 3.3V voltage output to power the entire CD player.

[0032] The present invention has the following advantages: the user's voice signal is collected by calling the microphone of the AI ​​voice recognition module, and the collected voice signal is digitized and processed into a digital audio signal; the digital audio signal is preprocessed, and the preprocessing includes noise reduction, filtering, and framing; the preprocessed digital audio signal is subjected to voice signal feature parameter extraction, and the voice signal feature parameter includes a Mel-frequency cepstral coefficient feature vector, and the acoustic characteristics of the voice signal are characterized by the voice signal feature parameter; the extracted voice signal feature parameter is input into a pre-trained voice recognition model based on a deep learning algorithm, and the input voice signal feature parameter is classified and recognized by the voice recognition model to obtain a text control instruction for the CD player. The present invention can accurately understand the user's voice instruction, automatically perform playback control and music recommendation according to the user's personalized needs and preferences, and meet the user's pursuit of convenient, intelligent, and personalized audio experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the drawings required for the implementation methods or the prior art descriptions are briefly introduced below. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0034] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0035] Figure 1 A schematic diagram of the flow of the AI ​​voice interactive CD playback control method provided in an embodiment of the present invention;

[0036] Figure 2 A schematic diagram of the speech recognition process in the AI ​​speech interactive CD playback control method provided in an embodiment of the present invention;

[0037] Figure 3 A schematic diagram of speech signal processing by a speech recognition model in the AI ​​speech interactive CD playback control method provided in an embodiment of the present invention;

[0038] Figure 4 A schematic diagram of the hardware architecture of an AI voice interactive CD playback control device provided in an embodiment of the present invention;

[0039] Figure 5 This is a schematic diagram of the power supply module in the AI ​​voice interactive CD playback control device provided in an embodiment of the present invention.

[0040] In the figure, 1. Main control module; 2. AI voice recognition module; 3. CD player audio link module; 4. Audio playback module; 5. Storage module; 6. Voice recognition chip; 7. Microphone; 8. Movement; 9. Driver chip; 10. Audio decoding chip; 11. Dual-mode wireless module; 12. Mobile server module; 13. Display module; 14. TFT touch color screen; 15. Power module; 16. Battery; 17. 3.3V voltage regulator chip; 18. Charging management chip; 19. Charging interface. DETAILED DESCRIPTION

[0041] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] See also Figure 1 and Figure 2 , an embodiment of the present invention provides an AI voice interactive CD playback control method, comprising the following steps:

[0043] S1. Call the microphone of the AI speech recognition module to collect the user's speech signal, and digitize the collected speech signal into a digital audio signal;

[0044] S2. Preprocess the digital audio signal, and the preprocessing includes noise reduction, filtering, and frame segmentation;

[0045] S3. Extract speech signal feature parameters from the preprocessed digital audio signal. The speech signal feature parameters include Mel frequency cepstral coefficient feature vectors, and the acoustic features of the speech signal are characterized by the speech signal feature parameters;

[0046] S4. Input the extracted speech signal feature parameters into a pre-trained speech recognition model constructed based on a deep learning algorithm, and classify and recognize the input speech signal feature parameters through the speech recognition model to obtain a text control instruction for the CD player device.

[0047] In this embodiment, in step S1, the AI speech recognition module collects the user's speech signal through the microphone, digitizes the speech signal, and converts it into a digital audio signal. Then in step S2, the digital audio signal is preprocessed, including operations such as noise reduction, filtering, and frame segmentation using the noise reduction processor FM1288 chip to improve the quality and recognizability of the speech signal. Among them, in step S3, the feature parameters of the speech signal are extracted, such as feature vectors such as Mel frequency cepstral coefficients (MFCC), and these feature vectors can effectively characterize the acoustic features of the speech signal. Furthermore, in step S4, the extracted feature vectors are input into a pre-trained speech recognition model. The speech recognition model is constructed based on a deep learning algorithm (such as a deep neural network or a convolutional neural network or a recurrent neural network, etc.), and after training with a large amount of speech data, it can classify and recognize the input feature vectors and output the corresponding text instructions.

[0048] In this embodiment, in step S2, during the preprocessing of the digital audio signal:

[0049] Define an array to store the sampling values of the digital audio signal;

[0050] The array format is: int audioSamples[BUFFER_SIZE], where BUFFER_SIZE represents the size of the array for caching audio sampling values;

[0051] Take the sampling value array of the digital audio signal and the array size as parameters, and calculate the average value of the sampling values in each window in units of a set window size by looping through the array;

[0052] A nested loop is constructed, wherein the nested loop includes an outer loop and an inner loop; the sliding position of the window is controlled by the outer loop, and the sampling values ​​in the window are accumulated and averaged by the inner loop; and the sampling value at the center of the window is replaced by the average value.

[0053] Specifically, the denoising method for preprocessing digital audio signals adopts a mean filtering algorithm. First, an array is defined in the software program to store the sample values ​​of the audio signal, such as int audioSamples[BUFFER_SIZE]; BUFFER_SIZE indicates that the array size of the cached audio sample values ​​can be set according to actual conditions. The function receives the array of audio sample values ​​and the array size as parameters. Inside the function, by looping through the array, the average value of the sample values ​​in each window is calculated in units of the set window size. Nested loops can be used, with the outer loop controlling the sliding position of the window, and the inner loop used to accumulate the sample values ​​in the window and average them, and then the average value replaces the sample value at the center of the window.

[0054] In this embodiment, the outer loop starts from half of the window size and ends at the end of the array minus half of the window size, so that each sampling point that needs filtering is located at the center of the window;

[0055] The window range of the inner loop is centered on the current outer loop position i, and is extended by windowSize / 2 sampling points on the left and right, where windowSize is the window size of the mean filter.

[0056] Specifically, the inner loop accumulates the sample values ​​in the window according to the current window position determined by the outer loop. After the accumulation is completed, the average value is calculated, and this average value represents the average level of the sample values ​​in the window. The calculated average value replaces the original sample value at the center of the window, thereby achieving noise reduction processing for the sample value at this position. After performing such processing on each sample value in the array that may be at the center of the window, the mean filtering noise reduction of the entire audio signal is completed.

[0057] For example, if the window size is set to 5, when the outer loop reaches the third position of the array, the inner loop accumulates and averages the sample values ​​at the first to fifth positions, and then assigns the average value to the sample value at the third position, and so on, to complete the processing of the entire array.

[0058] The implementation code of the nested loop is as follows:

[0059] " / *meanFilter mean filter function; audioSamples[] array storing audio signal sampling values; bufferSize audio signal sampling value array size windowSize mean filter window size* /

[0060] void meanFilter(int audioSamples[],int bufferSize,int windowSize){

[0061] / / The outer loop is used to control the position of the window sliding on the audio sample value array, starting from half the window size

[0062] / / End at the end of the array minus half the window size. This ensures that each sampling point that needs to be filtered can be calculated at the center of the window.

[0063] for(int i = windowSize / 2; i <bufferSize-windowSize / 2;i++){

[0064]

[0065] See also Figure 3 In this embodiment, in step S3, in the speech recognition model constructed based on the deep learning algorithm, the output values ​​of the speech input signal are multiplied by the corresponding weights and then added together, and the obtained results are added with the bias value corresponding to the speech output signal, and finally the final output result of the speech output signal is obtained through the activation function.

[0066] The expression of the speech recognition model is:

[0067]

[0068] Where y(n) is the speech signal output by the model for the nth class, and w i is the weight of speech output signal i, x i is the speech input signal i, b i is the bias value corresponding to the speech output signal i.

[0069] Specifically, when the voice input signal enters the speech recognition model, each voice input signal has its corresponding weight. The speech recognition model will multiply the output value of the voice input signal by the corresponding weight. Then, add the output values ​​after multiplying the weights. The result of the addition is added with the bias value corresponding to the voice output signal. Finally, the sum obtained by the above calculation is input into the activation function. The activation function will perform a nonlinear transformation on this sum to obtain the final output result of the voice output signal. For example, assuming there are three voice input signals, the output values ​​are A, B, and C, the corresponding weights are a, b, and c, and the bias value is d, then Aa+Bb+C*c is calculated first, and then d is added, and this result is input into the activation function to obtain the final output result. This calculation process enables the speech recognition model to learn and capture the complex features and patterns in the voice signal, thereby achieving accurate speech recognition.

[0070] See also Figure 4 , an embodiment of the present invention further provides an AI voice interactive CD playback control device, which adopts an AI voice interactive CD playback control method of the above embodiment, including a main control module 1, an AI voice recognition module 2, a CD player audio link module 3, an audio playback module 4 and a storage module 5;

[0071] The AI ​​voice recognition module 2 includes a voice recognition chip 6 and a microphone 7, the microphone 7 is electrically connected to the voice recognition chip 6, the voice recognition chip 6 is electrically connected to the main control module 1, and the AI ​​voice recognition module 2 is used to collect the user's voice control instructions;

[0072] The CD player audio link module 3 includes a movement 8, a drive chip 9 and an audio decoding chip 10; a connection is established between the movement 8 and the drive chip 9, a connection is established between the drive chip 9 and the audio decoding chip 10, and a connection is established between the audio decoding chip 10 and the main control module 1; the movement 8 is used to scan the CD disc information through a laser head; the drive chip 9 is used to read audio data; the audio decoding chip 10 is used to decode the read audio data into an audio signal;

[0073] A connection is established between the audio playback module 4 and the main control module 1. The audio playback module 4 is provided with an audio power amplifier chip. The audio playback module 4 is used to play the original audio signal through the audio power amplifier chip and the configured speaker;

[0074] The storage module 5 is connected to the main control module 1 , and the storage module 5 is used to randomly store the audio signal converted by decoding the audio data.

[0075] In this embodiment, a dual-mode wireless module 11 is also included, and a connection is established between the dual-mode wireless module 11 and the main control module 1. The dual-mode wireless module 11 is used for wireless data transmission and control of the CD player. In addition, a mobile server module 12 is also included, and a connection is established between the dual-mode wireless module 11 and the main control module 1. The mobile server module 12 is used by the user to send a playback control instruction to the CD player through a mobile terminal.

[0076] Specifically, the dual-mode wireless module 11 is connected to the main control module 1, realizing the wireless data transmission and control function of the CD player. In terms of data transmission, the dual-mode wireless module 11 can transmit various data, such as audio files, device status information, etc., between the CD player and external devices (such as mobile phones, computers, etc.). In terms of control, the external device can send control instructions to the main control module 1 with the help of the dual-mode wireless module 11, thereby controlling the operation of the CD player.

[0077] The mobile server module 12 establishes a connection with the main control module 1 through the dual-mode wireless module 11. When the user operates on the mobile terminal (such as a mobile phone), the mobile terminal sends a play control instruction to the mobile server module 12. After the mobile server module 12 receives and processes these instructions, it transmits the instructions to the main control module 1 through the dual-mode wireless module 11. The main control module 1 performs specific operations on the CD player according to the received instructions, such as playing, pausing, switching tracks, etc. For example, when the user uses a mobile phone to connect to the CD player through the dual-mode wireless module 11, he can operate on the mobile phone to select the CD track to be played, and the dual-mode wireless module 11 transmits this instruction to the main control module 1, and the main control module 1 controls the CD player to perform the operation of playing the corresponding track.

[0078] When performing voice recognition interaction, the device microphone 7 collects voice, converts it into digital signals and extracts features. The main control module 1 coordinates and processes the data signal, and sends the feature data to the mobile server module 12 through the dual-mode wireless module 11. The mobile server module 12 compares and matches the trained model, obtains the recognition result, and then returns it to the main control module 1. After the main control module 1 analyzes, it performs operations or displays according to the results to complete the interaction.

[0079] In this embodiment, a display module 13 is further included. The display module 13 is configured with a TFT touch color screen 14. The display module 13 is connected to the main control module 1 and is used to display music lyrics data corresponding to the played audio signal.

[0080] Specifically, the display module 13 establishes a connection with the main control module 1. When the CD player plays an audio signal, the main control module 1 obtains the corresponding music lyrics data and transmits the data to the display module 13. After the TFT touch color screen configured by the display module 13 receives the lyrics data from the main control module 1, it displays the lyrics in a clear and colorful form through its internal display drive circuit and pixel matrix. If the user touches the color screen, the touch signal will be transmitted to the main control module 1, and the main control module 1 controls the display module 13 to update the displayed lyrics information according to the touch operation (such as turning pages, switching display content, etc.). For example, in the process of playing a song, the TFT touch color screen of the display module 13 displays the sentence or paragraph of lyrics currently being played in real time, which is convenient for users to sing along or view.

[0081] In this embodiment, a power module 15 is also included, and the power module 15 is configured with a battery 16, a 3.3V voltage regulator chip, a charging management chip 18 and a charging interface 19;

[0082] The battery 16 is used to power the entire CD player; the charging interface 19 is used to connect an external power adapter to charge the battery 16; the charging management chip 18 is used to regulate the charging process; the 3.3V voltage regulator chip 17 is used to convert the voltage input from the battery 16 or the charging interface 19 into a 3.3V voltage output.

[0083] See also Figure 5 Specifically, the battery 16 serves as the power source for the entire CD player, providing the required electrical energy for each component of the device to enable it to operate normally. When the battery 16 is low on power, an external power adapter is connected through the charging interface 19. After the current from the external power source is input, the charging management chip 18 starts working. The charging management chip 18 monitors and regulates the charging process, such as controlling the charging current and voltage, monitoring the power of the battery 16, preventing overcharging, etc., to ensure that the battery 16 can be charged safely and efficiently. During the operation of the device, the voltage output by the battery 16 or the voltage input through the charging interface 19 may not be a stable 3.3V. At this time, the 3.3V voltage regulator chip 17 comes into play, converting and stabilizing the input voltage, and outputting a voltage of 3.3V, providing a stable and reliable power supply for components that require 3.3V power supply.

[0084] In a possible embodiment, the audio decoding chip 10 of the AI ​​voice recognition module 2 adopts Silan Micro series SC6137D, SC6135B, SC9659P / Lingyang SPHE8104GW; the driver chip 9 of the CD player audio link module 3 adopts Silan Micro SA1466, SA1461 / SL0311 / Roma chip, and the movement 8 adopts M93BG6 / DA11; the charging management chip 18 adopts TP4054 / 4056 / 4057 / 5000; the radio chip adopts RDA5807MP.

[0085] In a possible embodiment, the audio power amplifier chip adopts SiPT series XPT4863 / HAA9809S / ESMT series / CS43L21; battery 16: 18650, 21700 lithium battery 16 / polymer lithium battery 16; the voice recognition chip 6 adopts the offline voice chip Jerry JL701 series / Jerry AC695 series / Mountain View series AP8064 / Rockchip series RK2108 / Al module designed based on Jerry series AC7911B8 / BA; the charging interface 19 can be a Type-C interface or other DC socket, and its input voltage can be 5V / 9V / 12V power supply; the storage module 5 adopts WINDAND25Q series; the dual-mode wireless module 11 uses the Wi-Fi+BT integrated module of the ESP32 series chip or AP6255 chip, which complies with the 802.11b / g / n / ac standard, supports 2.4G and 5G dual-bands, the Wi-Fi interface is SDIO, and the Bluetooth interface is UART serial port.

[0086] In summary, the present invention calls the microphone 7 of the AI ​​speech recognition module 2 to collect the user's speech signal, and digitizes the collected speech signal into a digital audio signal; preprocesses the digital audio signal, and the preprocessing includes noise reduction, filtering, and framing; extracts speech signal feature parameters from the preprocessed digital audio signal, and the speech signal feature parameters include Mel frequency cepstral coefficient feature vectors, and the acoustic features of the speech signal are characterized by the speech signal feature parameters; the extracted speech signal feature parameters are input into a pre-trained speech recognition model based on a deep learning algorithm, and the input speech signal feature parameters are classified and recognized by the speech recognition model to obtain text control instructions for the CD player. When the speech input signal enters the speech recognition model, each speech input signal has its corresponding weight. The speech recognition model multiplies the output value of the speech input signal by its corresponding weight. Then, these output values ​​after multiplying the weights are added. The result of the addition is added with the bias value corresponding to the speech output signal. Finally, the sum obtained by the above calculation is input into the activation function. The activation function performs a nonlinear transformation on this sum, thereby obtaining the final output result of the speech output signal. The dual-mode wireless module 11 is connected to the main control module 1, realizing the wireless data transmission and control function of the CD player. In terms of data transmission, the dual-mode wireless module 11 can transmit various data, such as audio files, device status information, etc., between the CD player and external devices (such as mobile phones, computers, etc.). In terms of control, the external device can send control instructions to the main control module 1 with the help of the dual-mode wireless module 11, thereby controlling the operation of the CD player. The mobile server module 12 establishes a connection with the main control module 1 through the dual-mode wireless module 11. When the user operates on the mobile terminal (such as a mobile phone), the mobile terminal sends a playback control instruction to the mobile server module 12. After the mobile server module 12 receives and processes these instructions, it passes the instructions to the main control module 1 through the dual-mode wireless module 11. The main control module 1 performs specific operations on the CD player according to the received instructions, such as playing, pausing, switching tracks, etc. For example, when a user uses a mobile phone to connect to a CD player through the dual-mode wireless module 11, he can select the CD track to be played on the mobile phone, and the dual-mode wireless module 11 transmits this instruction to the main control module 1, and the main control module 1 controls the CD player to play the corresponding track. The display module 13 establishes a connection with the main control module 1. When the CD player plays an audio signal, the main control module 1 obtains the corresponding music lyrics data and transmits the data to the display module 13. After the TFT touch color screen configured by the display module 13 receives the lyrics data from the main control module 1, it displays the lyrics in a clear and colorful form through its internal display drive circuit and pixel matrix.If the user touches the color screen, the touch signal will be transmitted to the main control module 1, and the main control module 1 will control the display module 13 to update the displayed lyrics information according to the touch operation (such as turning pages, switching display content, etc.). For example, in the process of playing a song, the TFT touch color screen of the display module 13 displays the sentence or paragraph of lyrics currently being played in real time, which is convenient for users to sing along or view. The battery 16 serves as the power source for the entire CD player and provides the required power for each component of the device so that it can operate normally. When the battery 16 is low on power, an external power adapter is connected through the charging interface 19. After the current of the external power supply is input, the charging management chip 18 starts working. The charging management chip 18 monitors and regulates the charging process, such as controlling the charging current and voltage, monitoring the power of the battery 16, preventing overcharging, etc., to ensure that the battery 16 can be charged safely and efficiently. During the operation of the device, the voltage output by the battery 16 or the voltage input through the charging interface 19 may not be a stable 3.3V. At this time, the 3.3V voltage regulator chip plays a role, converting and stabilizing the input voltage, outputting a 3.3V voltage, and providing a stable and reliable power supply for components that require 3.3V power supply. The present invention can accurately understand the user's voice commands, automatically perform playback control and music recommendations according to the user's personalized needs and preferences, and meet the user's pursuit of convenient, intelligent, and personalized audio experience.

[0087] The present invention is described in more detail and in greater detail above through general description and specific embodiments. It should be understood that, based on the technical concept of the present invention, several conventional adjustments or further innovations can be made to these specific embodiments; however, as long as they do not deviate from the technical concept of the present invention, the technical solutions obtained by these conventional adjustments or further innovations also fall within the scope of protection of the claims of the present invention.

Claims

1. An AI voice interactive CD playback control method, characterized in that: The following steps are involved: The microphone of the AI ​​speech recognition module is called to collect the user's voice signal, and the collected voice signal is digitized and processed into a digital audio signal; Preprocessing the digital audio signal, wherein the preprocessing includes noise reduction, filtering, and framing; Extracting speech signal feature parameters from the preprocessed digital audio signal, wherein the speech signal feature parameters include Mel-frequency cepstral coefficient feature vectors, and representing the acoustic features of the speech signal through the speech signal feature parameters; The extracted voice signal feature parameters are input into a pre-trained voice recognition model constructed based on a deep learning algorithm, and the input voice signal feature parameters are classified and recognized by the voice recognition model to obtain text control instructions for the CD player.

2. The AI ​​voice interactive CD playback control method according to claim 1, characterized in that: During the preprocessing of the digital audio signal: Define an array to store the sample values ​​of the digital audio signal; The array format is: int audioSamples[BUFFER_SIZE], where BUFFER_SIZE represents the array size of the cached audio sample values; Taking the sampling value array and the array size of the digital audio signal as parameters, by looping through the array, taking the set window size as a unit, calculating the average value of the sampling values ​​in each window; Constructing a nested loop, wherein the nested loop includes an outer loop and an inner loop; controlling the sliding position of the window through the outer loop, and accumulating and averaging the sample values ​​in the window through the inner loop; Replace the sample value at the center of the window with the mean value.

3. The AI ​​voice interactive CD playback control method according to claim 2, characterized in that: The outer loop starts from half of the window size and ends at the end of the array minus half of the window size, so that each sampling point that needs filtering is located at the center of the window; The window range of the inner loop is centered on the current outer loop position i, and is extended by windowSize / 2 sampling points on the left and right. WindowSize is the window size of the mean filter.

4. The AI ​​voice interactive CD playback control method according to claim 1, characterized in that: In the speech recognition model constructed based on the deep learning algorithm, the output values ​​of the speech input signal are multiplied by the corresponding weights and then added together. The obtained results are added with the bias value corresponding to the speech output signal, and finally the final output result of the speech output signal is obtained through the activation function.

5. The AI ​​voice interactive CD playback control method according to claim 4, characterized in that: The expression of the speech recognition model is: Where y(n) is the speech signal output by the model for the nth class, and w i is the weight of speech output signal i, x i is the speech input signal i, b i is the bias value corresponding to the speech output signal i.

6. An AI voice interactive CD playback control device, using an AI voice interactive CD playback control method according to any one of claims 1 to 5, characterized in that: It includes a main control module (1), an AI voice recognition module (2), a CD player audio link module (3), an audio playback module (4) and a storage module (5); The AI ​​voice recognition module (2) comprises a voice recognition chip (6) and a microphone (7), the microphone (7) and the voice recognition chip (6) are electrically connected, the voice recognition chip (6) and the main control module (1) are electrically connected, the voice recognition chip (6) comprises an offline voice chip and an AI voice chip, and the AI ​​voice recognition module (2) is used to collect voice control instructions of a user; The CD player audio link module (3) comprises a movement (8), a drive chip (9) and an audio decoding chip (10); a connection is established between the movement (8) and the drive chip (9), a connection is established between the drive chip (9) and the audio decoding chip (10), and a connection is established between the audio decoding chip (10) and the main control module (1); the movement (8) is used to scan CD disc information through a laser head; the drive chip (9) is used to read audio data; and the audio decoding chip (10) is used to decode the read audio data and convert it into an audio signal; A connection is established between the audio playback module (4) and the main control module (1); the audio playback module (4) is provided with an audio power amplifier chip; and the audio playback module (4) is used to play the original audio signal through the audio power amplifier chip and a configured speaker; A connection is established between the storage module (5) and the main control module (1), and the storage module (5) is used to randomly store audio signals converted by decoding audio data.

7. The AI ​​voice interactive CD playback control device according to claim 6, characterized in that: It also comprises a dual-mode wireless module (11), a connection is established between the dual-mode wireless module (11) and the main control module (1), and the dual-mode wireless module (11) is used for wireless data transmission and control of the CD player device.

8. The AI ​​voice interactive CD playback control device according to claim 7, characterized in that: It also comprises a mobile server module (12), the mobile server module (12) establishing a connection with the main control module (1) via the dual-mode wireless module (11), and the mobile server module (12) is used for a user to send a playback control instruction to the CD playback device via a mobile terminal; The mobile server module (12) uses a trained speech recognition model constructed based on a deep learning algorithm for comparison and matching to obtain a recognition result, and transmits the recognition result back to the main control module (1). After the main control module (1) analyzes the recognition result, it performs an operation or displays it according to the analysis result to complete the interaction.

9. The AI ​​voice interactive CD playback control device according to claim 1, characterized in that: It also comprises a display module (13), the display module (13) being provided with a TFT touch color screen (14), the display module (13) being connected to the main control module (1), and the display module (13) being used to display music lyrics data corresponding to the played audio signal.

10. The AI ​​voice interactive CD playback control device according to claim 6, characterized in that: It also includes a power module (15), wherein the power module (15) is configured with a battery (16), a 3.3V voltage regulator chip (17), a charging management chip (18) and a charging interface (19); The battery (16) is used to supply power to the entire CD player; the charging interface (19) is used to connect an external power adapter to charge the battery (16); the charging management chip (18) is used to regulate the charging process; and the 3.3V voltage regulator chip (17) is used to convert the voltage input from the battery (16) or the charging interface (19) into a 3.3V voltage output to supply power to the entire CD player.

Citation Information

Cited By

  • CD playback control method and device based on ai speech interaction

    WO2026158137A1