Intelligent song listening and identifying method and system for gramophone

By collecting the original electrical signals from the phonograph cartridge and performing analog-to-digital conversion and intelligent processing, an equivalent model of the cartridge is constructed, enabling intelligent recognition and automatic repair of phonograph audio. This solves the problem of uncertainty in repair quality in existing technologies, improves recognition accuracy and repair quality, and enhances the user experience.

CN120954458APending Publication Date: 2025-11-14HIKEEN TECH SHENZHEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511091925.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing phonograph audio processing methods rely on manual identification and repair, lacking a precise acquisition path based on the original electrical signal of the phono cartridge and an intelligent music recognition and repair mechanism. This results in high uncertainty in repair quality and makes it unsuitable for large-scale digital archiving or automated content recognition needs.

Method used

By acquiring the raw electrical signal output from the phonograph cartridge, performing analog-to-digital conversion, extracting music feature vectors using a voiceprint recognition model, constructing an equivalent model of the cartridge, and running an adaptive filtering algorithm, frequency response equalization and distortion compensation are performed. Combined with style tag recognition and video resource matching, intelligent recognition and automatic repair are achieved.

Benefits of technology

It significantly improves the accuracy and restoration quality of phonograph audio recognition, provides an efficient and standardized digital audio recognition and restoration process, and enhances the user's immersive audiovisual experience and audio output selectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954458A_ABST
    Figure CN120954458A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent song listening and song identifying method and system for a gramophone, and the method comprises the steps: collecting an original electric signal outputted by a tone head of the gramophone, and carrying out the analog-to-digital conversion of the original electric signal, and obtaining a corresponding digital audio signal; extracting a music feature vector of the digital audio signal based on a preset voiceprint recognition model, and performing matching retrieval to obtain corresponding track information; under the condition that the track information identification fails, starting a local audio repairing process, constructing a cartridge equivalent model based on a preset frequency response characteristic modeling mechanism, and extracting a frequency domain response curve; and operating an adaptive filtering algorithm based on the frequency domain response curve, dynamically adjusting audio restoration parameters, executing frequency response equalization and distortion compensation operations, and generating restored audio data. The method and the device have the effect of improving the song identification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio signal processing, and in particular to a method and system for intelligent song recognition on a phonograph. Background Technology

[0002] Currently, in the process of audio recognition and digitization of sound sources from old-fashioned phonographs, the usual method is to collect the sound with a microphone or pickup, and then identify the music and improve the sound quality through manual timbre judgment or manual restoration. This method mainly relies on the operator's subjective identification of the record's content and the experience-based adjustments made to the frequency response and timbre in the later stages.

[0003] Common technical approaches in existing phonograph audio restoration workflows include: using a microphone to capture analog sound propagating in the air, manually operating a digital audio workstation (DAW) for filtering, gain adjustment, compression, and limiting, or using waveform observation to assist in audio restoration. However, in practical applications, this approach is limited by issues such as ambient noise interference, inconsistent human experience, and lack of standardized processing, resulting in significant uncertainty in restoration quality, low accuracy in music recognition, and an inability to meet the needs of large-scale digital archiving or automated content recognition.

[0004] The existing technical solutions mentioned above have the following drawbacks: the existing phonograph audio processing methods rely on manual identification and manual repair, lack a precise acquisition path based on the original electrical signal of the phono cartridge and an intelligent song recognition and repair processing mechanism, making it difficult to achieve an efficient and standardized digital audio recognition and restoration process, and therefore there is room for improvement. Summary of the Invention

[0005] To improve the accuracy of song recognition, this application provides a method and system for intelligent song recognition using a phonograph.

[0006] The above-mentioned objective of this application is achieved through the following technical solution: A method for intelligent song recognition on a phonograph, comprising: The original electrical signal output from the phonograph cartridge is acquired, and the original electrical signal is converted from analog to digital to obtain the corresponding digital audio signal; Based on a preset voiceprint recognition model, the music feature vector of the digital audio signal is extracted, and a matching retrieval is performed to obtain the corresponding track information; If the track information recognition fails, the local audio repair process is initiated, and an equivalent model of the phono cartridge is constructed based on the preset frequency response characteristic modeling mechanism, and the frequency domain response curve is extracted. Based on the frequency domain response curve, an adaptive filtering algorithm is run to dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data. Based on a predefined music style recognition model, style tag recognition is performed on the repaired audio data to obtain style tag information. Based on the style tag information, video resources corresponding to the style tag information are selected from a preset style matching database. The repaired audio data is compared with the original electrical signal for timbre similarity analysis. Based on the analysis results, a replica version of the audio that retains the original sound style and a lossless version of the audio that has been enhanced are output respectively.

[0007] By adopting the above technical solution, the original electrical signal output from the phonograph cartridge can be acquired and converted from analog to digital, thus accurately obtaining the original analog audio signal at the cartridge level. This avoids environmental noise and airborne distortion, significantly improving the signal quality foundation for subsequent identification and repair processes. By extracting music feature vectors based on a preset voiceprint recognition model and performing matching retrieval, the corresponding track information can be intelligently identified, improving the automatic identification efficiency of the phonograph source. When identification fails, a local audio repair process is initiated, constructing an equivalent model of the cartridge and extracting the frequency response curve. This simulates equipment characteristic errors and performs frequency response analysis, thereby achieving adaptive repair. The system provides pre-restored modeling support; by running an adaptive filtering algorithm to perform frequency response equalization and distortion compensation, it can automatically adjust audio restoration parameters, effectively eliminating frequency unevenness and harmonic distortion, thereby improving the fidelity and auditory consistency of the restored audio; by identifying style tags in the restored audio and matching them with corresponding video resources, it can enhance the style consistency of the final playback effect, thereby improving the user's immersive audiovisual experience; by performing timbre similarity analysis between the restored audio and the original signal and outputting a replica version or a lossless version based on the score, it can provide differentiated output solutions based on the degree of timbre preservation, thereby meeting the selection needs of different users between retro fidelity and high-quality listening experience.

[0008] In one example, this application can be further configured as follows: the acquisition of the original electrical signal output from the phonograph cartridge, and the analog-to-digital conversion of the original electrical signal to obtain the corresponding digital audio signal include: The analog signal output by the phono cartridge is acquired through an input structure, which includes a low-noise operational amplifier and a power isolation module to suppress ground potential fluctuations and radio frequency interference. The analog signal is sampled using an analog-to-digital converter chip, and a corresponding multi-channel digital audio stream is generated based on clock synchronization logic to obtain the corresponding digital audio signal.

[0009] By adopting the above technical solution, the analog signal output by the phono cartridge is acquired through the input structure, and interference is suppressed by using low-noise operational amplifiers and power isolation modules, which can reduce the influence of external electromagnetic interference and ground potential fluctuations, thereby ensuring signal sampling quality. By using analog-to-digital conversion chips and clock synchronization logic to generate multi-channel digital audio streams, the timing accuracy and channel consistency of audio signal digital processing can be guaranteed, thereby improving the stability of recognition and repair.

[0010] In one example, this application can be further configured as follows: the extraction of music feature vectors from the digital audio signal based on a preset voiceprint recognition model, and the matching retrieval to obtain the corresponding track information include: The digital audio signal is preprocessed by the audio feature extraction network in the preset voiceprint recognition model to extract multidimensional acoustic features, including Mel frequency cepstral coefficients, spectral centroid and rhythmic structure. Based on the multidimensional acoustic features, a music feature vector is generated, and the similarity between the music feature vector and the pre-stored standard feature vector is calculated to determine the target track information with the highest matching degree, thereby obtaining the track information.

[0011] By adopting the above technical solution, the audio feature extraction network in the voiceprint recognition model extracts multi-dimensional acoustic features such as Mel frequency cepstral coefficients, spectral centroid, and rhythmic structure, which can comprehensively capture the timbre and structural features of phonograph audio signals, thereby improving the comprehensiveness and expressiveness of feature matching. By calculating the similarity between the generated music feature vector and the standard feature vector, efficient identification of track information can be achieved, thereby improving the accuracy of track recognition.

[0012] In one example, this application can be further configured as follows: the construction of the phono cartridge equivalent model based on the preset frequency response characteristic modeling mechanism and the extraction of the frequency domain response curve include: By performing multi-band excitation tests on the phonograph cartridge output signal, amplitude and phase response data at different frequencies were collected. The equivalent circuit model of the phono cartridge is obtained by fitting the response data. The equivalent circuit model is constructed based on the RLC series resonant structure and is used to simulate the amplitude-frequency characteristics of the phono cartridge throughout the entire audio frequency band. Based on the equivalent circuit model, a fast Fourier transform is performed to extract the frequency domain response curve of the phono cartridge within the target frequency band. The frequency domain response curve is used as a reference for subsequent adaptive filtering and repair parameter adjustment.

[0013] By adopting the above technical solutions, multi-band excitation tests are conducted on the output signal of the phono cartridge, and amplitude and phase response data are collected. This allows for the establishment of a response characteristic model of the phono cartridge at different frequencies, thus providing a data foundation for subsequent repair strategies. By constructing an equivalent circuit model based on the RLC series resonant structure, the frequency response curve and resonance characteristics of a real phono cartridge can be simulated, thereby achieving simulation and reproduction of the frequency domain response. By using fast Fourier transform to extract the frequency domain response curve, efficient frequency domain data conversion and visualization analysis can be achieved, thus providing adjustable parameter basis for filtering and compensation algorithms.

[0014] In one example, this application can be further configured as follows: the adaptive filtering algorithm based on the frequency domain response curve, dynamically adjusting the audio restoration parameters, performing frequency response equalization and distortion compensation operations, and generating restored audio data includes: Based on the frequency domain response curve, a frequency compensation mapping table is constructed to identify the trough segment and nonlinear enhancement segment in the frequency response curve and assign the corresponding frequency band gain coefficient. The coefficients of the finite impulse response filter are updated in real time using the least mean square adaptive algorithm, and multi-band dynamic gain adjustment is performed to achieve equalization compensation in the frequency response imbalance region. Based on the harmonic distortion distribution characteristics estimated in the equivalent circuit model, a frequency domain pre-distortion signal is injected to cancel the final-stage nonlinear distortion, and finally the repaired audio data is generated.

[0015] By adopting the above technical solutions, and by constructing a frequency compensation mapping table and identifying the gain attenuation segment in the frequency response curve, the key frequency bands affecting sound quality can be identified, thus providing targeted targets for filter gain adjustment. By dynamically updating the coefficients of the finite impulse response filter through the LMS adaptive algorithm, the frequency band gain can be adjusted in real time to compensate for the problem of uneven frequency response, thereby achieving dynamic frequency response equalization. By estimating the harmonic distortion distribution based on the equivalent circuit model and injecting a pre-distortion signal, the final-stage nonlinear distortion can be canceled in advance, thereby improving the THD level of the audio output and enhancing the clarity of the sound quality.

[0016] In one example, this application can be further configured as follows: the step of filtering video resources corresponding to the style tag information from a preset style matching database based on the style tag information includes: The style tag information is converted into a style encoding to generate a standardized music style feature code, which is used to match the metadata index tag of each video resource in the database; Similarity retrieval is performed based on the aforementioned music style feature code, and multiple candidate video resources are selected from the style matching database; Based on the historical playback rating, content completeness and available resolution of each candidate video resource, a weighted score is calculated, and all candidate video resources are ranked by weighted score. The video resource with the highest score in the weighted scoring ranking is determined as the target MV resource, and the target MV resource and the repaired audio data are output synchronously.

[0017] By adopting the above technical solutions, and encoding style tag information into standardized musical style feature codes and matching them with the index tags of video resources, consistent mapping of style semantics can be achieved, thereby improving the accuracy of video matching. By filtering candidate video resources through similarity retrieval and weighting them based on playback rating, content completeness, and resolution, the optimal MV resource can be accurately selected, thereby improving the visual presentation quality of the playback content. By outputting the target MV resource synchronously with the repaired audio data, the user's auditory and visual collaborative experience can be enhanced, thereby improving the overall immersiveness of the playback.

[0018] In one example, this application can be further configured such that: the timbre similarity analysis between the repaired audio data and the original electrical signal includes: Phonic feature parameters are extracted from the repaired audio data and the original electrical signal, respectively. The timbre feature parameters include Mel frequency cepstral coefficients, spectral centroid, instantaneous energy, and zero crossover rate. The timbre feature parameters of the repaired audio data and the timbre feature parameters of the original electrical signal are compared using a cosine similarity function to generate a similarity score. Based on the similarity score, an output threshold is set to determine whether to retain the replica audio or output the enhanced audio.

[0019] By adopting the above technical solution, the timbre feature parameters of the repaired audio and the original signal are extracted separately and compared using the cosine similarity function. This allows for the quantification of the timbre similarity between different audio versions, thus providing an objective scoring basis for output decisions. By setting a similarity scoring threshold to determine whether to output a replica version or an enhanced version of the audio, a differentiated output strategy based on the degree of timbre preservation can be achieved, thereby satisfying users' diverse preferences in terms of original sound reproduction and auditory enhancement.

[0020] The second objective of this invention is achieved through the following technical solution: A phonograph intelligent music recognition system, comprising: The signal acquisition module is used to acquire the raw electrical signal output by the phonograph cartridge and perform analog-to-digital conversion on the raw electrical signal to obtain the corresponding digital audio signal. The matching module is used to extract the music feature vector of the digital audio signal based on a preset voiceprint recognition model, perform matching retrieval, and obtain the corresponding track information; The repair module is used to initiate a local audio repair process when track information recognition fails. It constructs an equivalent model of the phono cartridge based on a preset frequency response characteristic modeling mechanism and extracts the frequency domain response curve. The adjustment module is used to run an adaptive filtering algorithm based on the frequency domain response curve, dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data. The filtering module is used to identify style tags in the repaired audio data based on a predefined music style recognition model to obtain style tag information, and to filter video resources corresponding to the style tag information from a preset style matching database based on the style tag information. The analysis module is used to perform timbre similarity analysis between the repaired audio data and the original electrical signal, and output a replica version of the audio that retains the original sound style and a lossless version of the audio that has been enhanced, based on the analysis results.

[0021] By adopting the above technical solution and setting up signal acquisition, matching, repair, adjustment, filtering and analysis modules, the phonograph audio can be constructed into a complete process chain from analog acquisition to intelligent recognition, automatic repair and personalized output, thereby realizing a highly integrated, closed-loop intelligent phonograph music recognition system, improving the overall recognition efficiency and the level of intelligent sound quality repair.

[0022] In summary, this application includes the following beneficial technical effects: 1. By acquiring the raw electrical signal output from the phonograph cartridge and performing analog-to-digital conversion, the original analog audio signal at the cartridge level can be accurately obtained, avoiding environmental noise and airborne distortion, thus significantly improving the signal quality foundation for subsequent identification and repair processes; by extracting music feature vectors based on a preset voiceprint recognition model and performing matching retrieval, the corresponding track information can be intelligently identified, thereby improving the automatic identification efficiency of the phonograph sound source; by initiating a local audio repair process when identification fails, constructing an equivalent model of the cartridge and extracting the frequency domain response curve, the characteristic error of the device can be simulated and frequency response analysis can be performed, thus achieving modeling support before adaptive repair; 2. By running an adaptive filtering algorithm to perform frequency response equalization and distortion compensation, the audio restoration parameters can be automatically adjusted to effectively eliminate frequency unevenness and harmonic distortion, thereby improving the fidelity and auditory consistency of the restored audio. By identifying style tags and matching them with corresponding video resources after restoration, the style consistency of the final playback effect can be enhanced, thereby improving the user's immersive audiovisual experience. By performing timbre similarity analysis between the restored audio and the original signal and outputting a replica version or a lossless version based on the score, differentiated output solutions can be provided according to the degree of timbre preservation, thereby meeting the selection needs of different users between retro fidelity and high-quality listening experience. Attached Figure Description

[0023] Figure 1 This is a flowchart of a method for intelligent song recognition on a phonograph according to one embodiment of this application; Figure 2 This is a flowchart illustrating the implementation of step S10 in a phonograph intelligent song recognition method according to an embodiment of this application. Figure 3 This is a flowchart illustrating the implementation of step S20 in a phonograph intelligent song recognition method according to an embodiment of this application; Figure 4 This is a flowchart illustrating the implementation of step S30 in a phonograph intelligent song recognition method according to an embodiment of this application; Figure 5 This is a flowchart illustrating the implementation of step S40 in a phonograph intelligent song recognition method according to an embodiment of this application; Figure 6 This is a flowchart illustrating the implementation of step S50 in a phonograph intelligent song recognition method according to an embodiment of this application. Figure 7 This is a flowchart illustrating the implementation of step S60 in a phonograph intelligent song recognition method according to an embodiment of this application; Figure 8 This is a schematic diagram of a phonograph intelligent song recognition system according to one embodiment of this application. Detailed Implementation

[0024] The present application will be further described in detail below with reference to the accompanying drawings.

[0025] In one embodiment, such as Figure 1 As shown, this application discloses a method for intelligent song recognition on a phonograph, which specifically includes the following steps: S10: Collects the original electrical signal output from the phonograph cartridge and performs analog-to-digital conversion on the original electrical signal to obtain the corresponding digital audio signal.

[0026] Specifically, the weak electrical signal after the stylus and cartridge are connected is pre-buffered and amplified by an analog acquisition circuit, and the acquired analog audio signal is read in real time. Based on this, A / D conversion is performed. During the conversion, data sampling is performed using a timed interrupt method, and a sampling rate of 48kHz or higher is set to ensure the high-frequency response restoration accuracy. After the conversion is completed, the analog voltage value corresponding to each sampling cycle is converted into the corresponding discrete audio value and cached as a PCM format data block. The acquisition logic supports dual-channel input and timestamp marking to ensure the temporal consistency processing requirements of the subsequent audio processing module for the source data.

[0027] S20: Extract the music feature vector of the digital audio signal based on the preset voiceprint recognition model, perform matching retrieval, and obtain the corresponding track information.

[0028] Specifically, the sampled audio frame data is segmented and processed using a window function. A Fast Fourier Transform (FFT) is then performed on each frame of audio signal to extract the spectral feature information of each frame in the frequency domain and combine them into a spectral representation matrix of the overall audio sample. This spectrum is then input into a pre-trained voiceprint recognition model. The voiceprint recognition model adopts a multi-layer convolutional neural network structure. During the training phase, clearly labeled multi-source music from a public music library is used as training samples. A feature vector embedding space is constructed using a triplet loss function. The optimized model can stably extract music fingerprint features in key dimensions such as melody direction, rhythm pattern, and timbre envelope. After feature extraction, a cosine similarity comparison is performed with the feature vectors in the local music library. The comparison result with the highest similarity is selected as the candidate for recognition and preliminary music information is generated.

[0029] In one embodiment, the voiceprint recognition model is an audio fingerprint extraction model based on a deep convolutional neural network (CNN). It uses a publicly available music library and manually curated phonograph sample audio sources as the training dataset. The training data covers multiple record eras, pressing processes, and different genres and styles. Each audio sample undergoes standardized sampling rate conversion, dynamic range compression, and subjective label alignment before training. The model's input is a two-dimensional spectrogram tensor processed by short-time Fourier transform, and the output is a fixed-length 128-dimensional music embedding vector. During the training phase, a triplet loss function is used to optimize the model parameters. By minimizing the distance between different segments of the same track and maximizing the embedding distance between different tracks, effective separation and aggregation of the music feature space are achieved. The training process uses the Adam optimizer and is iteratively completed under a multi-GPU distributed architecture. The final generated voiceprint recognition model can perform low-dimensional feature encoding on any input digital audio data and supports efficient vector similarity retrieval.

[0030] S30: In the event of failure to recognize track information, initiate the local audio repair process, construct an equivalent model of the phono cartridge based on the preset frequency response characteristic modeling mechanism, and extract the frequency domain response curve.

[0031] Specifically, if the maximum similarity value returned by the recognition module is lower than the set recognition confidence threshold, the audio restoration process is triggered and the modeling stage begins. The system constructs multiple frequency excitation signals based on the spectral features embedded in the audio content to be restored, which are used to simulate frequency response analysis. Then, the output amplitude and phase information of each frequency under the excitation are collected to form a response sample set. The equivalent circuit model conforming to the RLC series structure is then constructed using the least squares fitting algorithm to simulate the gain characteristics of the phono cartridge in the frequency domain. This model is standardized into an input-response function form and frequency scanning is performed with a step interval of 100Hz. The amplitude and phase frequency curves of the output frequency between 20Hz and 20kHz are used as the frequency domain response curves, providing a basis for subsequent filter adjustment and audio enhancement.

[0032] In one embodiment, the phono cartridge frequency response modeling module constructs a parameterized equivalent model based on frequency scanning experiments, collects a large number of actual output data samples of different types of phono cartridges under full-frequency excitation, constructs a frequency-amplitude-phase triplet dataset, and uses the least squares fitting algorithm to fit the equivalent electrical parameters of the RLC series resonant structure model to establish a frequency response simulation function. Furthermore, combined with spectrum reconstruction and residual modeling methods, second-order and third-order harmonic amplitude response fitting curves are constructed at the resonant point and cutoff boundary to describe the harmonic gain and phase shift distribution in the nonlinear response region. The model supports calling pre-trained templates according to the cartridge model and also supports real-time reconstruction for accuracy enhancement. The final modeling result output is a frequency domain response curve and a harmonic gain spectrum, which are used by the subsequent filter parameter initialization, adaptive compensation control and distortion cancellation modules.

[0033] S40: Based on the frequency domain response curve, an adaptive filtering algorithm is run to dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data.

[0034] Specifically, the frequency domain response curve is segmented to identify frequency bands with significant amplitude attenuation. Gain correction factors are calculated based on these bands. The least mean square (LMS) adaptive algorithm is used to drive the weight adjustment process of the finite impulse response (FIR) filter kernel. In each iteration, real-time parameter tuning is performed based on the error gradient between the current frame and the desired target spectrum. Simultaneously, a predistortion mapping curve is constructed in parallel to compensate for harmonic anomalies caused by nonlinear distortion in the spectrum. This compensation signal is generated by model prediction and injected into the main filtering path for distortion cancellation. Finally, the processed frequency domain signal is inversely transformed back to the time domain to generate a continuous audio stream as the repaired audio data.

[0035] S50: Based on a predefined music style recognition model, style tag recognition is performed on the repaired audio data to obtain style tag information. Based on the style tag information, video resources corresponding to the style tag information are selected from a preset style matching database.

[0036] Specifically, the repaired audio data undergoes beat detection and frame-by-frame spectrogram analysis. Multiple music theory dimensions, including rhythm density, timbre brightness, and low-frequency energy distribution, are extracted to form a style feature vector. This feature vector is then input into a trained style recognition model for style classification prediction. The style recognition model is a deep neural network structure incorporating an attention mechanism. During training, multiple publicly available multi-style music libraries are used for multi-label supervised learning, and a label enhancement mechanism is introduced to improve the robustness of style boundary determination. After identifying the style label, the video resource retrieval logic is entered. Based on the feature label, index matching is performed with the style index field of each record in the preset MV video database. Text description similarity and historical preference scores are calculated, and after filtering and sorting, the best matching video is extracted as the style matching result.

[0037] In one embodiment, the music style recognition model is constructed based on a multi-label classification network with an attention fusion mechanism. It uses a labeled music library containing over 20 styles, including pop, rock, jazz, classical, and electronic, as the training set. During the preprocessing stage, rhythm density maps, energy change curves, tonality distribution, and high-order timbre vectors are extracted from the training data as input features. The model structure includes a three-layer convolutional feature extraction module and a one-layer multi-head attention aggregation layer, supporting concurrent output of multiple style labels. The training objective function is a binary cross-entropy loss superimposed with a style relevance sparse regularization term to improve the classification robustness for samples with ambiguous style boundaries. Model training uses K-fold cross-validation and a style distribution-balanced resampling strategy to ensure the model still has generalization ability under uneven sample distribution conditions. The final generated style recognition model can map audio of any length into an output containing style weight vectors, supporting subsequent style-driven content matching and recommendation tasks.

[0038] S60: Performs timbre similarity analysis between the repaired audio data and the original electrical signal, and outputs a replica audio version that retains the original sound style and an enhanced lossless audio version based on the analysis results.

[0039] Specifically, timbre feature sets such as MFCC, spectral centroid, zero-crossing rate, and temporal energy envelope are extracted from the original electrical signal and the restored audio data, respectively. Vector representations are established for each pair of audio segments and standardized processing is performed. The timbre comparison module is called to perform frame-by-frame comparison and sliding window aggregation using the cosine similarity function or Euclidean distance metric function to generate a global similarity score. If the score value exceeds the set fidelity threshold, it enters the replica version output channel. This version retains more of the original noise and frequency response pattern for style restoration display. If the score value does not meet the condition, it enters the enhancement processing channel. In this channel, multi-band dynamic range compression and enhancement control of human ear sensitive frequency bands are performed, and finally, a lossless version audio file with a more balanced signal-to-noise ratio and frequency response curve is output.

[0040] By adopting the above technical solution, the original electrical signal output from the phonograph cartridge can be acquired and converted from analog to digital, thus accurately obtaining the original analog audio signal at the cartridge level. This avoids environmental noise and airborne distortion, significantly improving the signal quality foundation for subsequent identification and repair processes. By extracting music feature vectors based on a preset voiceprint recognition model and performing matching retrieval, the corresponding track information can be intelligently identified, improving the automatic identification efficiency of the phonograph source. When identification fails, a local audio repair process is initiated, constructing an equivalent model of the cartridge and extracting the frequency response curve. This simulates equipment characteristic errors and performs frequency response analysis, thereby achieving adaptive repair. The system provides pre-restored modeling support; by running an adaptive filtering algorithm to perform frequency response equalization and distortion compensation, it can automatically adjust audio restoration parameters, effectively eliminating frequency unevenness and harmonic distortion, thereby improving the fidelity and auditory consistency of the restored audio; by identifying style tags in the restored audio and matching them with corresponding video resources, it can enhance the style consistency of the final playback effect, thereby improving the user's immersive audiovisual experience; by performing timbre similarity analysis between the restored audio and the original signal and outputting a replica version or a lossless version based on the score, it can provide differentiated output solutions based on the degree of timbre preservation, thereby meeting the selection needs of different users between retro fidelity and high-quality listening experience.

[0041] In one embodiment, such as Figure 2 As shown, in step S10, the original electrical signal output from the phonograph cartridge is acquired, and the original electrical signal is converted from analog to digital to obtain the corresponding digital audio signal. Specifically, this includes: S11: The analog signal output from the phono cartridge is acquired through the input structure, which includes a low-noise operational amplifier and a power isolation module to suppress ground potential fluctuations and radio frequency interference.

[0042] Specifically, the phono output port is connected to the signal input channel. A high-input-impedance differential amplifier connected in series at the input terminal linearly amplifies the millivolt-level audio signal generated by the phono. At the same time, an operational amplifier with low noise characteristics is used to maintain the signal-to-noise ratio of the audio signal during amplification. An isolated DC-DC converter module is connected before the power supply port of the operational amplifier to form an electrical isolation structure between the power supply path and the signal path to suppress noise coupling interference caused by ground potential fluctuations. Meanwhile, an RF filter network is configured in parallel before the operational amplifier to filter out high-frequency interference sources such as AM band and switching power supply. A symmetrical wiring structure is introduced in the entire analog signal chain to enhance the anti-common-mode interference capability. Finally, the noise-suppressed clean analog audio signal is output to the subsequent analog-to-digital converter module as a sampling input source.

[0043] S12: The analog signal is sampled using an analog-to-digital converter chip, and the corresponding multi-channel digital audio stream is generated based on the clock synchronization logic to obtain the corresponding digital audio signal.

[0044] Specifically, the differential input of the analog-to-digital converter chip is connected from the output of the low-noise operational amplifier. The analog audio signal is loaded into the internal sampling register of the analog-to-digital converter in a continuous stream manner. The high-speed sampling control logic is started to perform equally spaced sampling at a rate of not less than 96kHz. The sampling clock signal is uniformly generated by the main control unit and synchronized to each channel to ensure the consistency of the data frame timing between the two channels. The sampling result is converted into a fixed-width digital code value by the internal multi-bit voltage comparator. After the data conversion is completed, the output is a time-aligned alternating data stream of the left and right channels. The data is spliced ​​and the channel identifier is inserted using the interleaving buffer. Finally, a standard digital audio frame stream containing channel fields, timestamp fields and PCM data fields is generated. The continuous audio frame stream is stored in the audio buffer to be processed and used as the basic data input for subsequent voiceprint extraction and repair processing.

[0045] In one embodiment, such as Figure 3 As shown, in step S20, the music feature vector of the digital audio signal is extracted based on the preset voiceprint recognition model, and a matching retrieval is performed to obtain the corresponding track information. Specifically, this includes: S21: The digital audio signal is preprocessed by the audio feature extraction network in the preset voiceprint recognition model to extract multidimensional acoustic features, including Mel frequency cepstral coefficients, spectral centroid and rhythmic structure.

[0046] Specifically, the input digital audio signal undergoes DC drift removal and short-time frame segmentation, with each frame length set to 25 milliseconds and a frame shift set to 10 milliseconds. After weighting the frame shape using a Hamming window, a Fast Fourier Transform is performed to obtain frame-level spectral information. The spectrum is then weighted and integrated based on a pre-defined Mel filter bank to obtain an equivalent Mel frequency domain energy distribution map. Logarithmic operations and Discrete Cosine Transform are further performed on this energy map to extract Mel frequency cepstral coefficients to express timbre structure features. Simultaneously, the energy centroid position of each frame is calculated to form spectral centroid parameters to capture frequency band distribution characteristics. A rhythm map is constructed based on the inter-frame energy change rate to characterize the rhythm structure distribution. The three types of acoustic features extracted are combined in tensor form into a multi-dimensional feature matrix, which is then input into the subsequent network module of the recognition model for unified modeling processing.

[0047] S22: Generate music feature vectors based on multi-dimensional acoustic features, and calculate the similarity between the music feature vectors and pre-stored standard feature vectors to determine the target track information with the highest matching degree, thus obtaining the track information.

[0048] Specifically, the input multidimensional acoustic feature matrix is ​​fed into the loaded deep voiceprint coding network. Feature dimensionality reduction is performed according to the set order of each convolutional layer and pooling layer in the network to generate a high-dimensional embedding vector. The embedding vector is represented as a floating-point vector structure of uniform length to express the voiceprint fingerprint characteristics of the entire audio segment. Then, the similarity of this feature vector is compared with each reference music vector in the standard feature library. During the comparison, the cosine similarity measure is used to evaluate the angle similarity between audio features. Each round of comparison outputs a similarity score. All scores are sorted from high to low and the feature record with the highest similarity is selected as the matching result. The track information fields associated with the matching result include the song title, singer, recording version and style tag, etc., and are used as the final output of this recognition process.

[0049] In one embodiment, such as Figure 4 As shown, in step S30, the equivalent model of the phono cartridge is constructed based on the preset frequency response characteristic modeling mechanism, and the frequency domain response curve is extracted. Specifically, this includes: S31: By performing multi-band excitation tests on the output signal of the phonograph cartridge, amplitude and phase response data at different frequencies are collected.

[0050] Specifically, the excitation control logic connected to the phono cartridge input is configured with frequency scanning. The excitation frequency range is set to cover 20Hz to 20kHz and divided into at least 50 frequency points with logarithmic intervals as the test frequency sequence. For each frequency point, a sinusoidal signal with constant amplitude is output and continuously applied to the phono cartridge input side. The duration of each excitation is controlled between 300ms and 500ms to filter out transient distortion. The response signal is sampled synchronously with a fixed clock, and the output voltage waveform is acquired during the steady-state phase. The phase difference between the reference signal and the sampled signal is used to calculate the accurate phase response, and the amplitude response is extracted by combining the amplitude of the output envelope waveform. The gain (amplitude response) and phase shift (phase response) results corresponding to each frequency point are recorded respectively. Finally, the frequency-amplitude-phase triplet sequence is obtained as the frequency response modeling dataset. This dataset will be used as the parameter fitting input in the next stage of equivalent modeling.

[0051] S32: The equivalent circuit model of the phono cartridge is obtained by fitting the response data. The equivalent circuit model is based on the RLC series resonant structure and is used to simulate the amplitude-frequency characteristics of the phono cartridge throughout the entire audio frequency band.

[0052] Specifically, the frequency-amplitude-phase tripartite data obtained during the excitation test phase is imported into the frequency response modeling module. The least squares error fitting algorithm is used to reconstruct the amplitude-frequency curve of the measured data. The objective function is set as the minimum mean square error between the actual response value and the output value of the theoretical circuit model. Using an RLC series resonant circuit as the structural template, the three parameters of resistance, inductance, and capacitance are optimized and solved as variables. During the fitting process, the resonant peak position and bandwidth boundary are constrained to be consistent with the measured response to maintain physical authenticity. A set of equivalent circuit expressions that can be used for frequency domain simulation is constructed through the electrical parameter values ​​obtained by fitting. These expressions can be used to simulate the gain characteristics of the phono cartridge at various frequency points and support subsequent interpolation reconstruction and filter parameter derivation.

[0053] S33: Perform a fast Fourier transform based on the equivalent circuit model to extract the frequency domain response curve of the phono cartridge within the target frequency band. The frequency domain response curve is used as a reference for subsequent adaptive filtering and repair parameter adjustment.

[0054] Specifically, after completing the construction of the equivalent circuit model, the input signal is set as a unit impulse response, and time-domain response calculation is performed in the simulation environment. The response result is used as a time-domain sequence input to the Fast Fourier Transform function for transformation processing. The corresponding complex spectrum data is obtained, and the amplitude spectrum and phase spectrum curves are extracted as frequency-domain response outputs. If the target frequency band is focused between 1kHz and 16kHz, bandpass filtering is performed with this frequency band as a window, and interpolation enhancement is performed on the data points within the frequency band to improve resolution. Finally, the frequency and the corresponding amplitude curve are output as structured frequency-domain response curve data. This data serves as a reference for filter construction, dynamic gain adjustment, and compensation factor derivation, and participates in the core parameter setting process of the audio restoration module.

[0055] After the output frequency domain response curve is completed, the parameter tuning engine is invoked to perform a difference analysis between the current curve and the historical calibrated phono head response sample library. By calculating the average response difference across the entire frequency band and the residual distribution of key frequency bands, it is determined whether the current phono head has significant attenuation, abnormal peaks, or nonlinear shifts. If so, the dynamic strategy adjustment process is initiated. The initial gain factor setting ratio in the frequency compensation mapping table is automatically adjusted according to the difference magnitude. At the same time, the initial step size and learning rate parameters of the adaptive filter are updated to improve convergence efficiency. In the harmonic compensation path, the upper limit of distortion suppression amplitude and harmonic response weight are dynamically configured to balance the filter response range and repair intensity. During the adjustment process, the prior parameter template of a specific model of phono head can be loaded with reference to the device number or identification code to enhance adaptability. After the update is completed, the adjusted configuration parameters are synchronously transmitted to subsequent modules to realize the injection of processing strategies based on adaptive evolution of response anomaly characteristics.

[0056] In one embodiment, such as Figure 5 As shown, in step S40, an adaptive filtering algorithm is run based on the frequency domain response curve to dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data. Specifically, this includes: S41: Based on the frequency domain response curve, construct a frequency compensation mapping table, identify the trough segment and nonlinear enhancement segment in the frequency response curve, and assign the corresponding frequency band gain coefficient.

[0057] Specifically, using frequency domain response curve data as input, standardization and normalization are first performed to unify the response amplitudes of different frequency bands to the same comparison benchmark interval. After normalization, a sliding window method is used to calculate the local mean and variance characteristics of each frequency band. By setting an attenuation threshold, frequency band regions with continuous amplitudes significantly lower than the overall mean are identified as trough segments. At the same time, frequency band segments with sudden increases in amplitude accompanied by irregular phase changes are identified in the spectral morphology as nonlinear enhancement segments. Curvature analysis and first derivative changes are used to determine whether peak anomalies exceed the reasonable boundaries of the model. After identification, these key segments are mapped as frequency band labels on the frequency axis at a resolution of 10Hz or higher. Each labeled frequency band is assigned a corresponding target gain value, which is calculated based on the original response loss ratio or the nonlinear distortion inverse correction amount. Finally, a frequency compensation mapping table is constructed to drive the subsequent filter gain configuration and dynamic sound quality recovery process.

[0058] S42: The coefficients of the finite impulse response filter are updated in real time using the least mean square adaptive algorithm, and multi-band dynamic gain adjustment is performed to complete the equalization compensation for the frequency response imbalance region.

[0059] Specifically, the Finite Impulse Response (FIR) filter structure is initialized and a preset order is set. The generated frequency compensation mapping table is used as the reference for the filter gain objective function. Audio frame input is received in real time and filtering operation is performed. At the same time, the filtering output result is compared with the reference value of the target frequency response curve in the corresponding frequency band. The residual between the current output error and the reference target is calculated as an update factor and input into the Least Mean Square (LMS) algorithm core. The LMS algorithm controls the iterative adjustment direction and magnitude of the FIR filter weight vector according to the error and the learning rate parameter, so that the output error of each frame gradually converges to the minimum value. The real-time operation of the filter is not affected during the entire adjustment process, and it allows dynamic adaptation to the frequency response correction requirements as the input audio content changes. After the filter parameters are stable, the gain configuration can be further locked to achieve inter-band balance control, thereby improving the consistency and clarity of the overall audio output.

[0060] S43: Based on the harmonic distortion distribution characteristics estimated in the equivalent circuit model, a frequency domain pre-distortion signal is injected to cancel the final-stage nonlinear distortion, and finally, the repaired audio data is generated.

[0061] Specifically, based on the constructed RLC equivalent circuit model, the non-ideal performance of its amplitude-frequency characteristics in each frequency band is inferred. Combining the amplitude spectrum shape of the input signal, the theoretical distortion distribution map is calculated using the second and third harmonic estimation models. This map is used as the basis for constructing the frequency domain pre-distortion injection module. Subsequently, a set of pre-distortion spectrum vectors with opposite phase and reverse amplitude to the current audio spectrum structure is constructed as correction signals. The correction amplitude is calculated based on the harmonic type and nonlinear gain ratio of the frequency point. In the frequency domain space, the pre-distortion signal is superimposed and synthesized with the filter output spectrum. The synthesized spectrum is returned to the time domain by inverse Fourier transform to form the repaired audio frame data. After compensation, the audio distortion component can be suppressed to less than 20% of the original peak value at the THD (Total Harmonic Distortion) level. Subjectively, the purity of the timbre is improved and the detail recovery effect is enhanced. The final output repaired audio data will be used as the enhanced version audio under the recognition failure branch for subsequent style recognition and version output judgment.

[0062] In one embodiment, such as Figure 6 As shown, in step S50, which involves filtering video resources corresponding to the style tag information from a preset style matching database based on the style tag information, the specific steps include: S51: Perform style encoding conversion on style tag information to generate standardized music style feature codes, which are used to match the metadata index tags of each video resource in the database.

[0063] Specifically, the style tag information output by the audio style recognition module is input into the style encoding conversion logic. The fields such as style name, rhythm description, performance structure features, harmony type, and vocal style type contained in the tags are vector-encoded, and each style element is mapped to a set of semantic embedding representations in a multi-dimensional discrete vector space. By combining these embedding vectors and performing a unified standardization operation, a unique corresponding style feature code is generated. This style feature code uses a fixed-dimensional numerical vector expression and has good similarity calculation characteristics. At the same time, the feature code is aligned with the metadata tags of each video resource in the style matching database. Each video record in the database contains a style tag index field, a content key description field, and an encoding / decoding information field. The style tag field has been converted into an embedding form isomorphic to the style feature code in the preprocessing stage for subsequent similarity comparison and retrieval operations.

[0064] S52: Perform similarity retrieval based on music style feature codes and filter out multiple candidate video resources from the style matching database.

[0065] Specifically, the generated standardized style feature code is used as the query vector. The similarity calculation module is called to traverse and match the style index vectors of all video resources in the style matching database. During the matching process, cosine similarity is used as the main metric function. At the same time, a similarity lower limit threshold is set to filter low-relevance content. Each comparison outputs a score value and records the resource ID, matching tag and similarity score. All video resources that meet the threshold condition are selected as candidate sets. If the number of candidate results is insufficient, the threshold is dynamically relaxed for a second search. The total number of candidate video resources is finally controlled between 5 and 15 to facilitate subsequent weighted scoring and sorting operations and to ensure content diversity. After the selection is completed, the candidate list structure is cached in a temporary resource mapping table for use in the next stage of scoring.

[0066] S53: Calculate a weighted score for each candidate video resource based on its historical playback rating, content completeness, and available resolution, and then sort all candidate video resources by weighted score.

[0067] Specifically, the historical playback rating field, content integrity assessment flag, and video resolution field of each candidate video resource record are read sequentially. The historical playback rating is used as the main weighting factor and is included in the final scoring weighting model at a ratio of 0.5. The proportion of playable duration shown in the content integrity field and the data packet frame loss rate are inversely mapped to the integrity score and included in the total score calculation at a weight of 0.3. The available resolution information is converted into a high-definition rating level according to the pixel level and included in the model score at a weight of 0.2. After normalization, the three factors are summarized into a weighted total score for each candidate video. This score is used to reflect the comprehensive suitability of the target resource in terms of content quality, playback smoothness, and visual experience. Then, all candidate resources are sorted from high to low according to the weighted score to construct a video sorting list with a serial number index. This sorting result will be used for the final selection of the target video resource.

[0068] S54: Identify the video resource with the highest score in the weighted scoring ranking as the target MV resource, and output the target MV resource and the repaired audio data synchronously.

[0069] Specifically, the video resource with the highest score in the ranking results is selected, and its resource number, playback address, encoding and decoding information, and video track synchronization parameters are extracted into the target resource structure. The repaired audio data is used as the main channel content and bound to the target MV resource in a multimedia timeline. During the binding process, a timestamp interpolation alignment algorithm between the audio frame rate and the video frame rate is used to construct a synchronized timeline. A unified start point is set according to the audio beginning offset point and the video pre-frame position. At the same time, an audio buffer loading pre-trigger mechanism is set according to the video keyframe position to reduce playback delay. The video image and audio content are jointly packaged and pushed to the front-end output channel through the synchronization scheduling module, thereby achieving a consistent multimodal composite presentation output.

[0070] In one embodiment, such as Figure 7 As shown, in step S60, the repaired audio data and the original electrical signal are subjected to timbre similarity analysis, which specifically includes: S61: Extract timbre feature parameters from the restored audio data and the original electrical signal respectively. The timbre feature parameters include Mel frequency cepstral coefficients, spectral centroid, instantaneous energy, and zero crossover rate.

[0071] Specifically, the original electrical signal and the restored audio data are divided into frames, with each frame length set to 20 milliseconds and a frame shift set to 10 milliseconds. After smoothing the signal edges with a Hamming window function, a short-time Fourier transform is performed on each frame to obtain a time-frequency domain representation. Based on the spectrum, the Mel frequency cepstral coefficients are calculated for each frame to characterize its timbre envelope structure. The spectral centroid value is calculated by solving the centroid position of the power spectral density distribution to quantize the frequency energy shift. Instantaneous energy features are calculated based on the frame-level energy sum to reflect loudness variation characteristics. At the same time, the number of zero-crossing points in each frame is counted to form a zero-crossing rate parameter to represent signal roughness and the proportion of high-frequency components. The four types of feature parameters are cached in floating-point vector form, and complete timbre feature parameter matrices are constructed for both the original audio and the restored audio, providing a feature vector input basis for the subsequent similarity comparison process.

[0072] S62: The timbre feature parameters of the repaired audio data are compared with the timbre feature parameters of the original electrical signal using a cosine similarity function to generate a similarity score.

[0073] Specifically, the two sets of timbre feature matrices generated are paired and matched in frame time order. The feature vectors of each pair of corresponding frames are normalized and then cosine similarity is calculated. The specific calculation formula is the vector dot product divided by the product of their respective moduli to obtain the angle similarity value between the timbre parameters of each frame. The similarity scores of all frames are combined into a similarity sequence and a moving average filter is applied to smooth the effect of drastic fluctuations between frames. Finally, the mean of the entire sequence is aggregated to generate a global timbre similarity score. The score ranges from 0 to 1. The higher the value, the closer the repaired audio is to the original signal in terms of timbre characteristics. This score forms the core indicator for evaluating the fidelity of the repaired audio and will serve as the basis for judging the version output strategy.

[0074] S63: Set an output threshold based on the similarity score to determine whether to retain the replica audio or output the enhanced audio.

[0075] Specifically, the calculated timbre similarity score is compared with a preset timbre fidelity threshold. If the similarity score is higher than the threshold, the restored audio is considered to have basically preserved the timbre structure and auditory characteristics of the original electrical signal. In this case, the restored audio is directly output as a replica version to restore the original sound style and is available for use in fidelity-priority scenarios. If the similarity score is lower than the threshold, the audio enhancement process is triggered, and the dynamic range control algorithm is called to compensate and enhance the high-frequency weakened area. After high-frequency excitation and mid-frequency clarity enhancement processing are performed through human ear sensitive frequency band mapping, it is re-encoded into an enhanced version of the audio output. Finally, based on the score judgment result, a dynamic selection is made between conservative fidelity output and subjective auditory optimization output, realizing flexible switching between style reproduction and perceptual optimization.

[0076] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0077] In one embodiment, a phonograph intelligent song recognition system is provided, which corresponds one-to-one with the phonograph intelligent song recognition method described in the above embodiments. For example... Figure 8 As shown, this intelligent phonograph music recognition system includes a signal acquisition module, a matching module, a repair module, an adjustment module, a filtering module, and an analysis module. Detailed descriptions of each functional module are as follows: The signal acquisition module is used to acquire the raw electrical signal output from the phonograph cartridge and perform analog-to-digital conversion on the raw electrical signal to obtain the corresponding digital audio signal. The matching module is used to extract the music feature vector of the digital audio signal based on the preset voiceprint recognition model, perform matching retrieval, and obtain the corresponding track information; The repair module is used to initiate a local audio repair process when track information recognition fails. It constructs an equivalent model of the phono cartridge based on a preset frequency response characteristic modeling mechanism and extracts the frequency domain response curve. The adjustment module is used to run an adaptive filtering algorithm based on the frequency domain response curve, dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data. The filtering module is used to identify style tags in the repaired audio data based on a predefined music style recognition model, obtain style tag information, and filter video resources corresponding to the style tag information from a preset style matching database based on the style tag information. The analysis module is used to perform timbre similarity analysis between the repaired audio data and the original electrical signal. Based on the analysis results, it outputs a replica audio version that retains the original sound style and a lossless audio version that has been enhanced.

[0078] Optionally, the signal acquisition module includes: The acquisition submodule is used to acquire the analog signal output by the phono cartridge through the input structure, which includes a low-noise operational amplifier and a power isolation module to suppress ground potential fluctuations and radio frequency interference. The sampling submodule is used to sample analog signals using an analog-to-digital converter chip and generate corresponding multi-channel digital audio streams based on clock synchronization logic to obtain the corresponding digital audio signals.

[0079] Optional, the matching module includes: The preprocessing submodule is used to preprocess digital audio signals through the audio feature extraction network in the preset voiceprint recognition model to extract multidimensional acoustic features, including Mel frequency cepstral coefficients, spectral centroid and rhythmic structure. The similarity calculation submodule is used to generate music feature vectors based on multi-dimensional acoustic features, and to calculate the similarity between the music feature vectors and pre-stored standard feature vectors to determine the target track information with the highest matching degree, thus obtaining the track information.

[0080] Optional, the repair module includes: The test submodule is used to perform multi-band excitation tests on the output signal of the phonograph cartridge and collect amplitude and phase response data at different frequencies. The model building submodule is used to fit the equivalent circuit model of the phono cartridge based on the response data. The equivalent circuit model is built based on the RLC series resonant structure and is used to simulate the amplitude-frequency characteristics of the phono cartridge in the entire audio frequency band. The output curve submodule is used to perform a fast Fourier transform based on the equivalent circuit model to extract the frequency domain response curve of the phono cartridge within the target frequency band. The frequency domain response curve is used as a reference for subsequent adaptive filtering and repair parameter adjustment.

[0081] Optional, the adjustment modules include: The mapping table submodule is used to construct a frequency compensation mapping table based on the frequency domain response curve, identify the trough segment and nonlinear enhancement segment in the frequency response curve, and assign the corresponding frequency band gain coefficient. The equalization compensation submodule is used to update the coefficients of the finite impulse response filter in real time using the least mean square adaptive algorithm, perform multi-band dynamic gain adjustment, and complete the equalization compensation for the frequency response imbalance region. The audio repair submodule is used to inject a frequency domain pre-distortion signal based on the harmonic distortion distribution characteristics estimated in the equivalent circuit model, to cancel the final-stage nonlinear distortion, and finally generate the repaired audio data.

[0082] Optionally, the filtering module includes: The transcoding submodule is used to perform style encoding conversion on style tag information and generate standardized music style feature codes, which are used to match the metadata index tags of each video resource in the database. The retrieval submodule is used to perform similarity retrieval based on music style feature codes and filter out multiple candidate video resources from the style matching database. The sorting submodule is used to calculate a weighted score value based on the historical playback score, content completeness and available resolution of each candidate video resource, and sort all candidate video resources by weighted score. The audio output submodule is used to identify the video resource with the highest score in the weighted scoring ranking as the target MV resource, and to output the target MV resource and the repaired audio data synchronously.

[0083] Optionally, the analysis module includes: The extraction submodule is used to extract timbre feature parameters from the repaired audio data and the original electrical signal, respectively. The timbre feature parameters include Mel frequency cepstral coefficients, spectral centroid, instantaneous energy, and zero crossover rate. The comparison submodule is used to perform vector comparison between the timbre feature parameters of the repaired audio data and the timbre feature parameters of the original electrical signal using the cosine similarity function, and generate a similarity score. The threshold setting submodule is used to set the output threshold based on the similarity score to determine whether to retain the replica version audio or output the enhanced version audio.

[0084] For specific limitations regarding the intelligent song recognition system for phonographs, please refer to the limitations of the intelligent song recognition method for phonographs described above, which will not be repeated here. Each module in the aforementioned intelligent song recognition system for phonographs can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0085] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for intelligent song recognition on a phonograph, characterized in that, The aforementioned intelligent song recognition method for phonographs includes: The original electrical signal output from the phonograph cartridge is acquired, and the original electrical signal is converted from analog to digital to obtain the corresponding digital audio signal; Based on a preset voiceprint recognition model, the music feature vector of the digital audio signal is extracted, and a matching retrieval is performed to obtain the corresponding track information; If the track information recognition fails, the local audio repair process is initiated, and an equivalent model of the phono cartridge is constructed based on the preset frequency response characteristic modeling mechanism, and the frequency domain response curve is extracted. Based on the frequency domain response curve, an adaptive filtering algorithm is run to dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data. Based on a predefined music style recognition model, style tag recognition is performed on the repaired audio data to obtain style tag information. Based on the style tag information, video resources corresponding to the style tag information are selected from a preset style matching database. The repaired audio data is compared with the original electrical signal for timbre similarity analysis. Based on the analysis results, a replica version of the audio that retains the original sound style and a lossless version of the audio that has been enhanced are output respectively.

2. The method for intelligent song recognition on a phonograph according to claim 1, characterized in that, The process of acquiring the raw electrical signal output from the phonograph cartridge and performing analog-to-digital conversion on the raw electrical signal to obtain the corresponding digital audio signal includes: The analog signal output by the phono cartridge is acquired through an input structure, which includes a low-noise operational amplifier and a power isolation module to suppress ground potential fluctuations and radio frequency interference. The analog signal is sampled using an analog-to-digital converter chip, and a corresponding multi-channel digital audio stream is generated based on clock synchronization logic to obtain the corresponding digital audio signal.

3. The intelligent song recognition method for a phonograph according to claim 1, characterized in that, The process of extracting the music feature vector from the digital audio signal based on a preset voiceprint recognition model, performing matching retrieval, and obtaining the corresponding track information includes: The digital audio signal is preprocessed by the audio feature extraction network in the preset voiceprint recognition model to extract multidimensional acoustic features, including Mel frequency cepstral coefficients, spectral centroid and rhythmic structure. Based on the multidimensional acoustic features, a music feature vector is generated, and the similarity between the music feature vector and the pre-stored standard feature vector is calculated to determine the target track information with the highest matching degree, thereby obtaining the track information.

4. The intelligent song recognition method for a phonograph according to claim 1, characterized in that, The process of constructing a phono cartridge equivalent model based on a preset frequency response characteristic modeling mechanism and extracting the frequency domain response curve includes: By performing multi-band excitation tests on the phonograph cartridge output signal, amplitude and phase response data at different frequencies were collected. The equivalent circuit model of the phono cartridge is obtained by fitting the response data. The equivalent circuit model is constructed based on the RLC series resonant structure and is used to simulate the amplitude-frequency characteristics of the phono cartridge throughout the entire audio frequency band. Based on the equivalent circuit model, a fast Fourier transform is performed to extract the frequency domain response curve of the phono cartridge within the target frequency band. The frequency domain response curve is used as a reference for subsequent adaptive filtering and repair parameter adjustment.

5. The intelligent song recognition method for a phonograph according to claim 1, characterized in that, The process of running an adaptive filtering algorithm based on the frequency domain response curve, dynamically adjusting audio restoration parameters, performing frequency response equalization and distortion compensation operations, and generating restored audio data includes: Based on the frequency domain response curve, a frequency compensation mapping table is constructed to identify the trough segment and nonlinear enhancement segment in the frequency response curve and assign the corresponding frequency band gain coefficient. The coefficients of the finite impulse response filter are updated in real time using the least mean square adaptive algorithm, and multi-band dynamic gain adjustment is performed to achieve equalization compensation in the frequency response imbalance region. Based on the harmonic distortion distribution characteristics estimated in the equivalent circuit model, a frequency domain pre-distortion signal is injected to cancel the final-stage nonlinear distortion, and finally the repaired audio data is generated.

6. The intelligent song recognition method for a phonograph according to claim 1, characterized in that, The step of filtering video resources corresponding to the style tag information from a preset style matching database based on the style tag information includes: The style tag information is converted into a style encoding to generate a standardized music style feature code, which is used to match the metadata index tag of each video resource in the database; Similarity retrieval is performed based on the aforementioned music style feature code, and multiple candidate video resources are selected from the style matching database; Based on the historical playback rating, content completeness and available resolution of each candidate video resource, a weighted score is calculated, and all candidate video resources are ranked by weighted score. The video resource with the highest score in the weighted scoring ranking is determined as the target MV resource, and the target MV resource and the repaired audio data are output synchronously.

7. The intelligent song recognition method for a phonograph according to claim 1, characterized in that, The step of performing timbre similarity analysis between the repaired audio data and the original electrical signal includes: Phonic feature parameters are extracted from the repaired audio data and the original electrical signal, respectively. The timbre feature parameters include Mel frequency cepstral coefficients, spectral centroid, instantaneous energy, and zero crossover rate. The timbre feature parameters of the repaired audio data and the timbre feature parameters of the original electrical signal are compared using a cosine similarity function to generate a similarity score. Based on the similarity score, an output threshold is set to determine whether to retain the replica audio or output the enhanced audio.

8. A phonograph intelligent song recognition system, characterized in that, The aforementioned intelligent phonograph music recognition system includes: The signal acquisition module is used to acquire the raw electrical signal output by the phonograph cartridge and perform analog-to-digital conversion on the raw electrical signal to obtain the corresponding digital audio signal. The matching module is used to extract the music feature vector of the digital audio signal based on a preset voiceprint recognition model, perform matching retrieval, and obtain the corresponding track information; The repair module is used to initiate a local audio repair process when track information recognition fails. It constructs an equivalent model of the phono cartridge based on a preset frequency response characteristic modeling mechanism and extracts the frequency domain response curve. The adjustment module is used to run an adaptive filtering algorithm based on the frequency domain response curve, dynamically adjust the audio restoration parameters, perform frequency response equalization and distortion compensation operations, and generate restored audio data. The filtering module is used to identify style tags in the repaired audio data based on a predefined music style recognition model to obtain style tag information, and to filter video resources corresponding to the style tag information from a preset style matching database based on the style tag information. The analysis module is used to perform timbre similarity analysis between the repaired audio data and the original electrical signal, and output a replica version of the audio that retains the original sound style and a lossless version of the audio that has been enhanced, based on the analysis results.