An audio analysis method and system for differential localization
By employing a three-level hierarchical audio difference analysis method, including physical feature comparison, semantic deviation analysis, and pitch deviation analysis, the problem of heavy computational burden and noise interference in traditional audio difference analysis is solved. This method achieves high-precision and rapid audio difference localization and feedback, making it suitable for music teaching and audio quality monitoring.
Patent Information
- Application Number
- CN202411739569.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Traditional audio difference analysis methods are computationally intensive, slow in analysis speed, and susceptible to noise interference, resulting in insufficient accuracy and timeliness in audio difference localization.
A three-level hierarchical audio difference analysis method is adopted, including audio physical feature comparison, semantic deviation analysis and pitch deviation analysis. The signal is collected in real time by audio sensors and compared layer by layer. The semantic feature extraction model and pitch deviation analysis are used to gradually screen and accurately locate audio differences.
It significantly improves the analysis accuracy and noise resistance of audio difference localization, enabling faster response to audio difference detection needs and achieving accurate and rapid audio difference localization and feedback. It is suitable for scenarios such as music teaching and audio quality monitoring.
Smart Images

Figure CN119541546B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing, and more particularly to an audio analysis method and system for differential localization. Background Technology
[0002] Audio difference analysis refers to the process of comparing the characteristics of two audio signals to detect and locate the differences between them in order to identify deviations or anomalies. For example, in music teaching, real-time audio difference analysis can help students identify pitch, rhythm, and harmony deviations in their performances, and provide real-time feedback on incorrect notes, rhythmic deviations, or inaccurate intervals, helping students correct problems in their performances and improve their practice effectiveness.
[0003] In the field of audio difference analysis, traditional methods usually rely on comparing and analyzing the global high-dimensional features of audio signals, involving multiple dimensions such as frequency, timing, loudness, and pitch. While these methods can identify audio differences to a certain extent, they also bring problems such as heavy computational burden and slow analysis speed. Especially in high-noise environments, the robustness of the algorithm is significantly reduced. This method not only requires high-performance computing resources, but also often suffers from high data processing complexity and noise interference, resulting in the accuracy and timeliness of audio difference localization failing to meet the needs of real-time applications. Summary of the Invention
[0004] This invention addresses the technical problems of traditional audio difference analysis methods, such as heavy computational burden, slow analysis speed, and susceptibility to noise interference, which lead to insufficient accuracy and timeliness in audio difference localization. It provides an audio analysis method and system for difference localization to solve these problems.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0006] In a first aspect, the present invention provides an audio analysis method for differential localization, applied to an audio analysis system for differential localization, the system including a user terminal, the system being communicatively connected to an audio sensor, the audio sensor being detachably mounted on an audio source device, comprising: extracting audio physical features in response to a first audio signal uploaded by the user terminal to obtain reference audio feature time-series information; receiving a second audio signal collected by the audio sensor to extract audio physical features to obtain monitoring audio feature time-series information; comparing the reference audio feature time-series information and the monitoring audio feature time-series information with audio physical features to obtain a first indiscriminate audio moment, wherein the first indiscriminate audio moment has a first reference... The system generates an audio signal tag and a first monitoring audio signal tag; performs semantic deviation analysis on the first reference audio signal tag and the first monitoring audio signal tag to obtain a second indifferent audio moment, wherein the second indifferent audio moment has a second reference audio signal tag and a second monitoring audio signal tag; performs pitch deviation analysis on the second reference audio signal tag and the second monitoring audio signal tag to obtain a pitch difference audio moment, wherein the pitch difference audio moment has a third reference audio signal tag and a third monitoring audio signal tag; and adds the pitch difference audio moment, the third reference audio signal tag, and the third monitoring audio signal tag to the audio analysis result and sends it to the user terminal.
[0007] Secondly, the present invention provides an audio analysis system for differential localization, the system including a user terminal, the system being communicatively connected to an audio sensor, the audio sensor being detachably installed on an audio source device, and comprising: a reference audio feature timing information acquisition module, used to extract audio physical features in response to a first audio signal uploaded by the user terminal to obtain reference audio feature timing information; a monitoring audio feature timing information acquisition module, used to receive a second audio signal collected by the audio sensor to extract audio physical features to obtain monitoring audio feature timing information; and an audio physical feature comparison module, used to compare the reference audio feature timing information and the monitoring audio feature timing information to obtain a first indiscriminate audio moment, wherein the first indiscriminate audio moment has a first reference audio signal. The system includes a first reference audio signal tag and a first monitoring audio signal tag; a semantic deviation analysis module, used to perform semantic deviation analysis on the first reference audio signal tag and the first monitoring audio signal tag to obtain a second indifferent audio moment, wherein the second indifferent audio moment has a second reference audio signal tag and a second monitoring audio signal tag; a pitch deviation analysis module, used to perform pitch deviation analysis on the second reference audio signal tag and the second monitoring audio signal tag to obtain a pitch difference audio moment, wherein the pitch difference audio moment has a third reference audio signal tag and a third monitoring audio signal tag; and an audio analysis result sending module, used to add the pitch difference audio moment, the third reference audio signal tag, and the third monitoring audio signal tag to the audio analysis result and send it to the user terminal.
[0008] Thirdly, the present invention also provides an electronic device, comprising:
[0009] At least one processor; a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of the first aspects above.
[0010] Fourthly, a computer-readable storage medium storing a computer program that, when executed, implements the steps of the method described in any one of the first aspects above.
[0011] The beneficial effects of this invention are as follows: Compared with traditional methods that perform high-dimensional feature comparison on the entire audio segment, this invention adopts a three-level hierarchical audio difference analysis method, which can perform physical feature comparison, semantic deviation analysis, and pitch deviation analysis on the audio signal layer by layer. It can gradually filter and accurately locate audio differences, thereby reducing unnecessary data processing burden, significantly improving the analysis accuracy and noise resistance of audio difference location, and responding to audio difference detection needs more quickly. It can achieve accurate and fast audio difference location and feedback, and better adapt to application scenarios that require real-time analysis and high-precision feedback, such as music teaching and audio quality monitoring. Attached Figure Description
[0012] Figure 1 A flowchart illustrating an audio analysis method for differential localization provided by the present invention;
[0013] Figure 2 A schematic diagram of the structure of an audio analysis system for differential localization provided by the present invention;
[0014] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention;
[0015] Figure 4 This is a schematic diagram of the structure of a computer-readable storage medium provided by the present invention.
[0016] The components represented by each number in the attached diagram are explained below:
[0017] The system includes a reference audio feature timing information acquisition module 11, a monitoring audio feature timing information acquisition module 12, an audio physical feature comparison module 13, a semantic deviation analysis module 14, a pitch deviation analysis module 15, an audio analysis result sending module 16, an electronic device 500, a memory 510, a processor 520, a first computer program 511, a computer-readable storage medium 600, and a second computer program 611. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0020] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0021] Example 1, as Figure 1 As shown, this embodiment of the invention provides an audio analysis method for differential localization, applied to an audio analysis system for differential localization. The system includes a user terminal, and the system is communicatively connected to an audio sensor. The audio sensor is detachably installed on an audio source device, including:
[0022] S100: In response to the first audio signal uploaded by the user, perform audio physical feature extraction to obtain the reference audio feature timing information.
[0023] Specifically, the method provided by this invention is implemented through an audio analysis system for differential localization. The system includes a user terminal for receiving and uploading a reference audio signal to establish reference audio feature data, and receiving audio difference analysis results from the system, including specific feedback information such as physical feature deviations, semantic deviations, and pitch deviations, to help the user quickly understand and locate the deviation details in the audio. The system is communicatively connected to an audio sensor, meaning it can transmit data and communicate in real-time with the audio sensor wirelessly or via wired means to acquire and analyze real-time audio signal data from the sensor. This connection allows the system to monitor changes in audio features in real time and promptly detect anomalies. The audio sensor is designed to be flexibly mounted and detachable to adapt to different sound source devices. Users can install the audio sensor on the specific device that needs to be monitored and change the device or position as needed. For example, the audio sensor can be installed on various devices, such as musical instruments, broadcasting equipment, or industrial machines, by means of clamping, magnetic attraction, or bracket fixation, to meet diverse application needs. For example, in music teaching, the audio sensor can be temporarily installed on instruments such as pianos, guitars, or violins to monitor the student's playing audio and compare it with the teacher's reference audio through the system. After the real-time audio data collected by the audio sensor is transmitted to the system, the system analyzes the characteristics of the playing notes, rhythms, etc., and feeds back the deviation to the student to help them quickly adjust their playing.
[0024] First, the audio analysis system receives the first audio signal uploaded by the user, i.e., the reference audio file, such as a standard performance audio file. Next, it extracts the audio physical features of the first audio signal, that is, it extracts the physical features of the first audio signal using audio processing algorithms to obtain the core physical features of the first audio signal. Existing algorithms can be used for audio feature extraction; for example, the Fourier transform algorithm is used to convert the audio signal from the time domain to the frequency domain to extract frequency components; the short-time energy analysis algorithm is used to calculate the energy changes of the audio signal within different time windows to extract volume features; and autocorrelation analysis or zero-crossing rate detection is used to extract the rhythm and duration features of the audio signal. The resulting first physical feature extraction results include frequency, loudness, duration, rhythm, and pitch. Frequency analysis can determine the pitch and harmony of the reference audio; loudness features can monitor abnormal volume fluctuations in comparison; rhythm and duration can identify rhythm deviations; and pitch can compare the accuracy of notes. Then, the first physical feature extraction results are arranged in a time series to form the reference audio feature time-series information. For example, in music teaching, the teacher uploads a standard piano performance audio. By extracting the frequency, rhythm, pitch, and loudness features of this performance, the system generates the timing information of the baseline audio features of this performance. Next, the system compares the student's performance (monitoring audio) with this baseline audio layer by layer to detect the student's performance deviation.
[0025] S200: Receives the second audio signal collected by the audio sensor, extracts audio physical features, and obtains the timing information of the monitored audio features.
[0026] Specifically, audio data is collected in real time by an audio sensor installed on the audio source device, and the collected second audio signal (monitoring audio) is sent to the audio analysis system via a wired or wireless communication connection. Then, the physical features of the second audio signal, including frequency, loudness, duration, rhythm, and pitch, are extracted using an audio processing algorithm to obtain the second physical feature extraction result. The method for extracting the second physical feature is the same as the method for extracting the first physical feature. The second physical feature extraction result is then arranged in chronological order to generate monitoring audio feature time-series information. For example, in a music teaching scenario, the audio sensor collects real-time audio of a student's performance and uploads it to the audio analysis system. By extracting physical features such as frequency, loudness, and pitch of the performance, the system generates monitoring audio feature time-series information. Next, the system compares this information with the reference audio features of a standard performance to identify deviations in notes or rhythms and provides the results back to the student, enabling them to correct errors in their performance in a timely manner.
[0027] S300: Compare the audio physical features of the reference audio feature timing information and the monitoring audio feature timing information to obtain a first indistinguishable audio moment, wherein the first indistinguishable audio moment has a first reference audio signal label and a first monitoring audio signal label.
[0028] In one embodiment, step S300 of this application further includes:
[0029] S310: Perform a global comparison of the audio physical features of the reference audio feature time series information and the monitoring audio feature time series information according to the audio physical feature comparison function to obtain the global deviation coefficient of the audio physical features.
[0030] In one embodiment, step S310 of this application further includes:
[0031] S311: Constructing an audio physical feature comparison function:
[0032] , , ,
[0033] , , ,in, Characterizes the temporal information of the baseline audio features. ~ Characterizes the temporal information of the Q-attribute baseline audio features. Characterize and monitor the temporal information of audio features. ~ Characterize Q attributes to monitor audio feature temporal information. ~ Characterizes the temporal information of the baseline audio features of the i-th attribute. ~ Let N represent the temporal information of the monitored audio features of the i-th attribute, N represent the temporal number of the baseline audio feature temporal information of the i-th attribute, and M represent the temporal number of the monitored audio feature temporal information of the i-th attribute. Characterization adjustment parameters, >1, Characterizing the deviation of the physical features of the i-th attribute. Global deviation coefficient characterizing the physical characteristics of audio.
[0034] Specifically, firstly, an audio physical feature comparison function is constructed, in which... Characterizes the temporal information of the baseline audio features; ~ The temporal information representing the Q attributes of the reference audio features is the temporal information of the multiple attributes (such as frequency, loudness, pitch, etc.) contained in the reference audio. Characterizes and monitors audio features and temporal information; ~ The monitoring of audio features is characterized by Q attributes, that is, the monitoring of the temporal information of multiple attributes (such as frequency, loudness, pitch, etc.) contained in the audio. ~ The i-th attribute represents the temporal information of the baseline audio feature, where the i-th attribute is any one of the Q attributes; ~ N represents the temporal information of the monitored audio feature of the i-th attribute; M represents the temporal number of the baseline audio feature temporal information of the i-th attribute, that is, the number of temporal data points recorded on the i-th attribute of the baseline audio; Characterization adjustment parameters, >1, where, The default value is 10, which is used to adjust the degree of influence of timing deviations. For example, in a music teaching scenario, if the system is used to identify rhythm deviations in a student's performance, adjusting this parameter... It can be appropriately increased to ensure that even minor timing deviations can be amplified and identified, especially in equipment noise monitoring scenarios. The value can be lowered to reduce the impact of timing deviations and focus more on deviations in frequency or loudness; by adjusting the parameters... Its flexible settings allow for appropriate adjustment of the impact of timing deviations in different application scenarios, thereby meeting different analytical accuracy requirements; Characterizing the deviation of the physical features of the i-th attribute. A global deviation coefficient characterizing the physical features of audio, used to comprehensively evaluate benchmark audio and monitor the overall deviation of audio across all physical features.
[0035] Then, based on the audio physical feature comparison function, a global comparison of the audio physical features is performed on the reference audio feature time series information and the monitoring audio feature time series information to calculate the global deviation coefficient of the audio physical features.
[0036] S320: When the global deviation coefficient of the audio physical feature is greater than or equal to the global deviation coefficient threshold, extract the first indistinguishable audio moment and the physical feature difference audio moment, wherein the first indistinguishable audio moment refers to the moment when each attribute audio physical feature of the reference audio feature time series information and the monitoring audio feature time series information is the same, and the physical feature difference audio moment refers to the moment when any attribute audio physical feature of the reference audio feature time series information and the monitoring audio feature time series information is different; S330: When the global deviation coefficient of the audio physical feature is less than the global deviation coefficient threshold, set all moments of the reference audio feature time series information and the monitoring audio feature time series information as the first indistinguishable audio moment.
[0037] Specifically, firstly, a global deviation coefficient threshold is configured to determine whether there are significant deviations in the overall physical characteristics of the reference audio and the monitoring audio. A reasonable global deviation coefficient threshold can be configured based on the application scenario and accuracy requirements. For example, in high-precision music teaching, the global deviation coefficient threshold can be set lower to identify subtle note and rhythm deviations. Since a small amount of random error may occur during the physical feature comparison process, if the global error is below the preset threshold, it is determined that the physical characteristics of the reference audio and the monitoring audio are generally consistent, and no further difference comparison is needed. If the global error exceeds the threshold, it indicates that there are significant differences between the reference audio and the monitoring audio, requiring detailed analysis and discrete extraction of specific moments of difference.
[0038] Next, the global deviation coefficient of the audio physical feature is judged according to the global deviation coefficient threshold. When the global deviation coefficient of the audio physical feature is greater than or equal to the global deviation coefficient threshold, the moment when each attribute audio physical feature of the reference audio feature time series information and the monitoring audio feature time series information is the same is set as the first indifference audio moment, that is, the time point when the reference audio feature time series information and the monitoring audio feature time series information are the same in each attribute audio physical feature; and the moment when any attribute audio physical feature of the reference audio feature time series information and the monitoring audio feature time series information is different is set as the physical feature difference audio moment, that is, as long as there is a deviation in any physical feature (such as frequency or loudness), the moment will be marked as the physical feature difference audio moment, thus obtaining the first indifference audio moment and the physical feature difference audio moment.
[0039] When the global deviation coefficient of the audio physical features is less than the global deviation coefficient threshold, the physical features of the reference audio and the monitored audio are generally consistent. Therefore, all moments of the time series information of the reference audio features and the time series information of the monitored audio features are set as the first indistinguishable audio moments. By calculating the global deviation coefficient of the audio physical features and setting a global deviation coefficient threshold for judgment, random errors can be quickly eliminated when judging global deviations. Fine-grained comparisons are only performed at the moments when actual deviations are detected, thereby quickly screening audio segments with no significant differences. Computational resources are concentrated on processing key difference moments, reducing the computational burden and significantly improving the efficiency and accuracy of audio difference analysis, thus providing users with more accurate and faster feedback.
[0040] S400: Perform semantic deviation analysis on the first reference audio signal label and the first monitoring audio signal label to obtain a second indifferent audio moment, wherein the second indifferent audio moment has a second reference audio signal label and a second monitoring audio signal label.
[0041] In one embodiment, step S400 of this application further includes:
[0042] S410: Based on the semantic feature extraction model, perform semantic feature extraction on the first reference audio signal label and the first monitoring audio signal label to obtain the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal.
[0043] In one embodiment, step S410 of this application further includes:
[0044] S411: Configure a semantic vector mapping space, which is used to map characters into unique multidimensional mathematical vectors; S412: Construct a semantic feature extraction loss function based on the semantic vector mapping space. ,in, The loss value is extracted to represent semantic features. Characterized by the adjustment coefficient, 0 < b < 1. Representing the k-th dimension of the identifier vector, S413: Using the audio source device model as a constraint, collect audio signal recording datasets and audio semantic identifier datasets of the same model device; S414: Using the audio semantic identifier dataset as supervision and the audio signal recording dataset as input, train the semantic feature extraction model based on the semantic vector mapping space and the semantic feature extraction loss function.
[0045] Specifically, constructing a semantic feature extraction model begins with configuring a semantic vector mapping space. This space maps characters to unique multi-dimensional mathematical vectors, i.e., configuring a high-dimensional vector space (typically between 50 and 300 dimensions, with the specific dimensions set according to application requirements) to carry the vector representation of each semantic unit. Specific mapping methods (such as word embedding or character embedding) map characters, words, or musical notes to unique positions in the vector space. For example, using audio semantic embedding, each semantic unit is mapped to a corresponding multi-dimensional vector. For instance, the character "do" can be mapped to the vector [0.12, 0.34, -0.45, ...], while "re" can be mapped to another unique vector. By configuring the semantic vector mapping space, abstract semantic features can be transformed into operable vector representations. This not only facilitates rapid matching in subsequent semantic comparisons but also ensures the uniqueness and discriminative power of semantic features, significantly improving the system's efficiency and accuracy in difference detection. This enables more precise extraction and comparison of semantic features in audio.
[0046] Next, a semantic feature extraction loss function is constructed based on the semantic vector mapping space. This loss function optimizes the semantic feature extraction model, enabling it to extract and represent semantic features more accurately, ensuring that the model generates accurate and highly discriminative vector representations in the mapping space. In the semantic feature extraction loss function, The semantic feature extraction loss value is used to measure the accuracy of the model in extracting semantic features in the mapping space. The greater the similarity, the smaller the corresponding loss value. The adjustment coefficient is characterized by a range of 0 < b < 1. The adjustment coefficient is used to balance the weight of the loss function on the bias of each dimension vector, so that the model can flexibly adjust the feature extraction intensity in different scenarios. Representing the k-th dimension of the identifier vector, The k-th dimension of the extracted vector is represented by H, which represents the mathematical dimension of the vector, i.e., the total number of dimensions of the vector. This loss function quantifies the differences between similar semantic vectors, ensuring that the greater the similarity, the smaller the loss, thereby guiding the model to generate more accurate semantic feature extraction vectors. At the same time, by introducing an adjustment coefficient to control the flexibility of the loss, it can further optimize the model's accuracy and stability in extracting semantic features, making it suitable for high-precision semantic comparison and deviation detection scenarios.
[0047] Further, the model of the audio source device (such as model number, series category, etc.) is obtained, and using the audio source device model as the search constraint feature, audio signal recording datasets and audio semantic identifier datasets of the same model are collected. Through this search process, device data unrelated to the target model can be excluded, ensuring that the dataset only contains audio signal data of the target model. The audio signal recording data and audio semantic identifier data are in one-to-one correspondence; that is, each audio signal recording has a corresponding audio semantic identifier used to accurately describe the characteristics and semantic content of the audio. The audio signal recording data includes physical characteristic information such as frequency, pitch, loudness, rhythm, and interval. The audio semantic identifier data refers to the semantic description or label of the audio signal recording data, indicating the specific content or meaning contained in the audio signal. These identifiers can be used to describe the state, notes, emotions, playing techniques, etc. of the audio signal, such as "C note," "G note," "glissando," and "tremolo."
[0048] A semantic feature extraction model is constructed based on a convolutional neural network. The model includes an input layer, convolutional layers, pooling layers, fully connected layers, and an output layer. The input data of the input layer is the audio signal, and the output data of the output layer is the semantic features of the audio signal. Multiple convolutional layers are used to extract low-level and high-level features from the spectrum. Each convolutional layer extracts local features of the audio, such as frequency patterns and pitch variations, through a convolution kernel. The pooling layer reduces the data size by downsampling, thereby reducing computational complexity while retaining the main features. The fully connected layer combines the previously extracted local features to generate a more globally meaningful feature vector.
[0049] Next, using audio signal recording data as input and audio semantic identifier data as supervision, the semantic feature extraction model is trained under supervision based on the semantic vector mapping space and the semantic feature extraction loss function, using the audio signal recording dataset and the audio semantic identifier dataset as training data. First, the audio signal recording dataset is converted into a format suitable for convolutional neural network processing (such as spectrograms or Mel spectrograms) and used as model input. Then, the audio semantic identifier dataset is mapped to the semantic vector space to generate high-dimensional semantic vectors as supervision signals, enabling the model to learn the mapping relationship between audio features and semantic identifiers. Finally, the preprocessed audio signal is input into the semantic feature extraction model, which then processes the signal through a series of convolutional layers, pooling layers, and... The fully connected layer progressively extracts the physical and semantic features of the audio. Then, the distance between the model's output semantic vector and the target semantic vector is calculated using a semantic feature extraction loss function. The loss value reflects the similarity between the model's prediction and the true semantic label, which is used to optimize the model. Then, backpropagation and gradient descent algorithms are used to update the model's parameters (including convolutional kernels and fully connected layer weights) based on the loss value. During training, the model continuously adjusts its parameters to gradually reduce the difference between the output semantic feature vector and the target semantic vector. Then, the model is trained iteratively multiple times on the entire dataset, gradually reducing the loss value to achieve high accuracy in semantic feature extraction until the expected convergence accuracy is met, resulting in a trained semantic feature extraction model.
[0050] Finally, the first reference audio signal label and the first monitoring audio signal label are input into the semantic feature extraction model for semantic feature extraction, outputting the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal. By constructing a semantic feature extraction model, semantic information can be effectively extracted from complex audio signals, while improving the efficiency and accuracy of semantic information extraction, thus providing support for semantic deviation comparison.
[0051] S420: When the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal are the same, the second indistinguishable audio moment is obtained; S430: When the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal are different, the semantically different audio moment is obtained.
[0052] Specifically, when the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal are the same, it indicates that the semantic content of the two audio segments at that moment is consistent, that is, there is no semantic difference between the reference audio and the monitoring audio during this period. Then, this moment is marked as the second indistinguishable audio moment and added to the second indistinguishable audio moment. When all moments are analyzed, the second indistinguishable audio moment is output, wherein the second indistinguishable audio moment has a second reference audio signal label and a second monitoring audio signal label. When the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal are different, it indicates that there is a difference in the semantic content of the two audio segments at that moment. Then, this moment is marked as the semantically different audio moment, and the semantically different audio moment is obtained.
[0053] S500: Perform pitch deviation analysis on the second reference audio signal tag and the second monitoring audio signal tag to obtain the pitch difference audio moment, wherein the pitch difference audio moment has a third reference audio signal tag and a third monitoring audio signal tag.
[0054] In one embodiment, step S500 of this application further includes:
[0055] S510: Extract the reference pitch and reference frequency of the second reference audio signal tag; S520: Extract the monitoring pitch and monitoring frequency of the second monitoring audio signal tag; S530: When the reference pitch and the monitoring pitch are different, or / and the reference frequency and the monitoring frequency are different, obtain the interval difference audio moment.
[0056] Specifically, firstly, the reference pitch and reference frequency at that moment are extracted from the second reference audio signal tag as the reference interval information for that moment. The reference pitch represents the tone of the reference audio, and the reference frequency represents the frequency of that tone. This information serves as a reference for interval comparison. Then, the monitoring pitch and monitoring frequency at that moment are extracted from the second monitoring audio signal tag. Next, the reference pitch is compared with the monitoring pitch, and the reference frequency is compared with the monitoring frequency. If the reference pitch and the monitoring pitch are different and / or the reference frequency and the monitoring frequency are different, then that moment is designated as an interval difference audio moment. The interval difference audio moment has a third reference audio signal tag and a third monitoring audio signal tag. By determining interval differences based on pitch and frequency, interval difference audio moments can be identified efficiently and accurately, providing support for audio difference localization.
[0057] S600: Add the interval difference audio time, the third reference audio signal label, and the third monitoring audio signal label to the audio analysis results and send them to the user terminal.
[0058] In one embodiment, step S600 of this application further includes:
[0059] S610: Obtain the physical feature difference audio moment of the audio physical feature comparison between the reference audio feature timing information and the monitoring audio feature timing information, wherein the physical feature difference audio moment has a fourth reference audio signal label and a fourth monitoring audio signal label; S620: Obtain the semantic difference audio moment of the semantic deviation analysis between the first reference audio signal label and the first monitoring audio signal label, wherein the semantic difference audio moment has a fifth reference audio signal label and a fifth monitoring audio signal label; S630: Construct a pitch deviation triplet array based on the pitch difference audio moment, the third reference audio signal label and the third monitoring audio signal label; construct a physical feature deviation triplet array based on the physical feature difference audio moment, the fourth reference audio signal label and the fourth monitoring audio signal label; construct a semantic deviation triplet array based on the semantic difference audio moment, the fifth reference audio signal label and the fifth monitoring audio signal label; S640: Add the pitch deviation triplet array, the physical feature deviation triplet array and the semantic deviation triplet array to the audio analysis result and send it to the user terminal.
[0060] Specifically, the moment when there is a difference in the audio physical features during the comparison of the reference audio feature timing information and the monitoring audio feature timing information is defined as the physical feature difference audio moment, thereby obtaining the physical feature difference audio moment. The physical feature difference audio moment has a fourth reference audio signal label and a fourth monitoring audio signal label. The moment when there is a difference in the semantic deviation analysis of the first reference audio signal label and the first monitoring audio signal label is defined as the semantic difference audio moment, wherein the semantic difference audio moment has a fifth reference audio signal label and a fifth monitoring audio signal label.
[0061] Then, based on the interval difference audio moments, the third reference audio signal label, and the third monitoring audio signal label, an interval deviation triplet is constructed. That is, for each interval difference audio moment, it is combined with the corresponding third reference audio signal label and the third monitoring audio signal label to form an interval deviation triplet. Based on the physical feature difference audio moments, the fourth reference audio signal label, and the fourth monitoring audio signal label, a physical feature deviation triplet is constructed. That is, for each physical feature difference audio moment, it is combined with the corresponding fourth reference audio signal label and the fourth monitoring audio signal label to form a physical feature deviation triplet. Based on the semantic difference audio moments, the fifth reference audio signal label, and the fifth monitoring audio signal label, a semantic deviation triplet is constructed. That is, for each semantic difference audio moment, it is combined with the corresponding fifth reference audio signal label and the fifth monitoring audio signal label to form a semantic deviation triplet.
[0062] Finally, the interval deviation triplet array, the physical feature deviation triplet array, and the semantic deviation triplet array are integrated into a complete audio analysis result, which is then sent to the user terminal. By sending the interval, physical feature, and semantic deviation triplet arrays as audio analysis results to the user terminal, detailed multi-dimensional difference analysis information can be provided, enabling the user to fully grasp the differences in the audio signal and thus make efficient decisions and adjustments.
[0063] The audio analysis method for differential localization provided in this invention has at least the following technical effects:
[0064] 1. By adopting a three-level hierarchical audio difference analysis method, physical feature comparison, semantic deviation analysis, and pitch deviation analysis of audio signals can be performed layer by layer. This allows for the gradual screening and precise location of audio differences, thereby reducing unnecessary data processing burdens and significantly improving the analysis accuracy and noise resistance of audio difference location. It can respond to audio difference detection needs more quickly, achieving accurate and rapid audio difference location and feedback, and better adapting to application scenarios that require real-time analysis and high-precision feedback, such as music teaching and audio quality monitoring.
[0065] 2. By calculating the global deviation coefficient of audio physical features and setting a threshold for the global deviation coefficient, random errors can be quickly eliminated when judging global deviation. Fine comparison is only performed when actual deviation is detected, thereby quickly screening audio segments with no significant differences. Computational resources are concentrated on processing key differences, reducing the computational burden and significantly improving the efficiency and accuracy of audio difference analysis, thus providing users with more accurate and faster feedback.
[0066] Example 2, as Figure 2As shown, based on the same inventive concept as the audio analysis method for differential localization provided in Embodiment 1, this embodiment of the invention also provides an audio analysis system for differential localization. The system includes a user terminal, and the system is communicatively connected to an audio sensor. The audio sensor is detachably mounted on an audio source device, and includes:
[0067] The reference audio feature timing information acquisition module 11 is used to extract audio physical features in response to the first audio signal uploaded by the user terminal to obtain reference audio feature timing information; the monitoring audio feature timing information acquisition module 12 is used to receive the second audio signal collected by the audio sensor and extract audio physical features to obtain monitoring audio feature timing information; the audio physical feature comparison module 13 is used to compare the reference audio feature timing information and the monitoring audio feature timing information to obtain a first indiscriminate audio moment, wherein the first indiscriminate audio moment has a first reference audio signal label and a first monitoring audio signal label; the semantic deviation analysis module 14 is used to analyze the first reference audio... A semantic deviation analysis is performed on the signal tag and the first monitored audio signal tag to obtain a second indifferent audio moment, wherein the second indifferent audio moment has a second reference audio signal tag and a second monitored audio signal tag; a pitch deviation analysis module 15 is used to perform pitch deviation analysis on the second reference audio signal tag and the second monitored audio signal tag to obtain a pitch difference audio moment, wherein the pitch difference audio moment has a third reference audio signal tag and a third monitored audio signal tag; an audio analysis result sending module 16 is used to add the pitch difference audio moment, the third reference audio signal tag and the third monitored audio signal tag into the audio analysis result and send it to the user terminal.
[0068] In one embodiment, the audio analysis system for differential localization further includes:
[0069] The system obtains the physical feature difference audio moments from the comparison of the audio physical features of the reference audio feature timing information and the monitoring audio feature timing information, wherein the physical feature difference audio moments have a fourth reference audio signal label and a fourth monitoring audio signal label; it also obtains the semantic difference audio moments from the semantic deviation analysis of the first reference audio signal label and the first monitoring audio signal label, wherein the semantic difference audio moments have a fifth reference audio signal label and a fifth monitoring audio signal label; it constructs a pitch deviation triplet array based on the pitch difference audio moments, the third reference audio signal label and the third monitoring audio signal label, a physical feature deviation triplet array based on the physical feature difference audio moments, the fourth reference audio signal label and the fourth monitoring audio signal label, and a semantic deviation triplet array based on the semantic difference audio moments, the fifth reference audio signal label and the fifth monitoring audio signal label; and it adds the pitch deviation triplet array, the physical feature deviation triplet array and the semantic deviation triplet array to the audio analysis results and sends them to the user terminal.
[0070] In one embodiment, the audio analysis system for differential localization further includes:
[0071] The audio physical feature time series information of the reference audio feature and the time series information of the monitored audio feature are globally compared according to the audio physical feature comparison function to obtain the global deviation coefficient of the audio physical feature. When the global deviation coefficient of the audio physical feature is greater than or equal to the global deviation coefficient threshold, the first indistinguishable audio moment and the physical feature difference audio moment are extracted. The first indistinguishable audio moment refers to the moment when every attribute audio physical feature of the reference audio feature time series information and the time series information of the monitored audio feature is the same, and the physical feature difference audio moment refers to the moment when any attribute audio physical feature of the reference audio feature time series information and the time series information of the monitored audio feature is different. When the global deviation coefficient of the audio physical feature is less than the global deviation coefficient threshold, all moments of the reference audio feature time series information and the time series information of the monitored audio feature are set as the first indistinguishable audio moment.
[0072] In one embodiment, the audio analysis system for differential localization further includes:
[0073] Based on the semantic feature extraction model, semantic features are extracted from the first reference audio signal label and the first monitoring audio signal label to obtain the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal. When the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal are the same, the second indistinguishable audio moment is obtained. When the semantic features of the first reference audio signal and the semantic features of the first monitoring audio signal are different, the semantically different audio moment is obtained.
[0074] In one embodiment, the audio analysis system for differential localization further includes:
[0075] Configure a semantic vector mapping space, which maps characters to unique multidimensional mathematical vectors; construct a semantic feature extraction loss function based on the semantic vector mapping space. ,in, The loss value is extracted to represent semantic features. Characterized by the adjustment coefficient, 0 < b < 1. Representing the k-th dimension of the identifier vector, The k-th dimension of the extracted vector is represented by H, which represents the dimension of the mathematical vector. The audio signal recording dataset and audio semantic identifier dataset of the same model of audio source device are collected as constraints. The semantic feature extraction model is trained based on the semantic vector mapping space and the semantic feature extraction loss function, with the audio semantic identifier dataset as supervision and the audio signal recording dataset as input.
[0076] In one embodiment, the audio analysis system for differential localization further includes:
[0077] Extract the reference pitch and reference frequency of the second reference audio signal tag; extract the monitoring pitch and monitoring frequency of the second monitoring audio signal tag; when the reference pitch and the monitoring pitch are different, or / and the reference frequency and the monitoring frequency are different, obtain the interval difference audio moment.
[0078] In one embodiment, the audio analysis system for differential localization further includes:
[0079] Construct an audio physical feature comparison function:
[0080] , , , , , ,in, Characterizes the temporal information of the baseline audio features. ~ Characterizes the temporal information of the Q-attribute baseline audio features. Characterize and monitor the temporal information of audio features. ~ Characterize Q attributes to monitor audio feature temporal information. ~ Characterizes the temporal information of the baseline audio features of the i-th attribute. ~ Let N represent the temporal information of the monitored audio features of the i-th attribute, N represent the temporal number of the baseline audio feature temporal information of the i-th attribute, and M represent the temporal number of the monitored audio feature temporal information of the i-th attribute. Characterization adjustment parameters, >1, Characterizing the deviation of the physical features of the i-th attribute. Global deviation coefficient characterizing the physical characteristics of audio.
[0081] Example 3, please refer to Figure 3 , Figure 3 This is a schematic diagram illustrating an embodiment of the electronic device provided in this invention. For example... Figure 3 As shown, this embodiment of the invention provides an electronic device 500, including a memory 510, a processor 520, and a first computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the first computer program 511, it performs the following steps: extracting audio physical features in response to a first audio signal uploaded by a user terminal to obtain reference audio feature timing information; receiving a second audio signal collected by an audio sensor to extract audio physical features to obtain monitoring audio feature timing information; comparing the reference audio feature timing information and the monitoring audio feature timing information to obtain a first indistinguishable audio moment, wherein the first indistinguishable audio moment has a first... A reference audio signal tag and a first monitoring audio signal tag are used; semantic deviation analysis is performed on the first reference audio signal tag and the first monitoring audio signal tag to obtain a second indifferent audio moment, wherein the second indifferent audio moment has a second reference audio signal tag and a second monitoring audio signal tag; pitch deviation analysis is performed on the second reference audio signal tag and the second monitoring audio signal tag to obtain a pitch difference audio moment, wherein the pitch difference audio moment has a third reference audio signal tag and a third monitoring audio signal tag; the pitch difference audio moment, the third reference audio signal tag, and the third monitoring audio signal tag are added to the audio analysis result and sent to the user terminal.
[0082] Example 4, please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided by an embodiment of the present invention. For example... Figure 4As shown, this embodiment provides a computer-readable storage medium 600, on which a second computer program 611 is stored. When the second computer program 611 is executed by a processor, it performs the following steps: extracting audio physical features in response to a first audio signal uploaded by a user terminal to obtain reference audio feature timing information; receiving a second audio signal collected by an audio sensor to extract audio physical features to obtain monitoring audio feature timing information; comparing the reference audio feature timing information and the monitoring audio feature timing information to obtain a first indistinguishable audio moment, wherein the first indistinguishable audio moment has a first reference audio signal tag and a first monitoring audio signal tag. Signal tags; semantic deviation analysis is performed on the first reference audio signal tag and the first monitoring audio signal tag to obtain a second indifferent audio moment, wherein the second indifferent audio moment has a second reference audio signal tag and a second monitoring audio signal tag; pitch deviation analysis is performed on the second reference audio signal tag and the second monitoring audio signal tag to obtain a pitch difference audio moment, wherein the pitch difference audio moment has a third reference audio signal tag and a third monitoring audio signal tag; the pitch difference audio moment, the third reference audio signal tag, and the third monitoring audio signal tag are added to the audio analysis result and sent to the user terminal.
[0083] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0084] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0086] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0087] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0088] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.
[0089] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for audio analysis for differential positioning, characterized in that, The application is applied to an audio analysis system for differential positioning, the system comprises a user terminal, the system is in communication connection with an audio sensor, the audio sensor is detachably installed on a sound source device, and comprises: extracting audio physical features in response to a first audio signal uploaded by the user terminal to obtain reference audio feature time sequence information; extracting audio physical features from a second audio signal collected by the audio sensor to obtain monitoring audio feature time sequence information; comparing the reference audio feature time sequence information and the monitoring audio feature time sequence information to obtain a first non-difference audio time, wherein the first non-difference audio time has a first reference audio signal label and a first monitoring audio signal label; performing semantic deviation analysis on the first reference audio signal label and the first monitoring audio signal label to obtain a second non-difference audio time, wherein the second non-difference audio time has a second reference audio signal label and a second monitoring audio signal label; performing interval deviation analysis on the second reference audio signal label and the second monitoring audio signal label to obtain an interval difference audio time, wherein the interval difference audio time has a third reference audio signal label and a third monitoring audio signal label; adding the interval difference audio time, the third reference audio signal label and the third monitoring audio signal label into an audio analysis result and sending the audio analysis result to the user terminal; wherein comparing the reference audio feature time sequence information and the monitoring audio feature time sequence information to obtain a first non-difference audio time comprises: performing global comparison of audio physical features on the reference audio feature time sequence information and the monitoring audio feature time sequence information according to an audio physical feature comparison function to obtain an audio physical feature global deviation coefficient; when the audio physical feature global deviation coefficient is greater than or equal to a global deviation coefficient threshold, extracting the first non-difference audio time and a physical feature difference audio time, wherein the first non-difference audio time refers to a time when each attribute audio physical feature of the reference audio feature time sequence information and the monitoring audio feature time sequence information is the same, and the physical feature difference audio time refers to a time when any one attribute audio physical feature of the reference audio feature time sequence information and the monitoring audio feature time sequence information is different; when the audio physical feature global deviation coefficient is less than the global deviation coefficient threshold, setting all times of the reference audio feature time sequence information and the monitoring audio feature time sequence information as the first non-difference audio time; wherein performing semantic deviation analysis on the first reference audio signal label and the first monitoring audio signal label to obtain a second non-difference audio time comprises: extracting semantic features from the first reference audio signal label and the first monitoring audio signal label according to a semantic feature extraction model to obtain first reference audio signal semantic features and first monitoring audio signal semantic features; when the first reference audio signal semantic features and the first monitoring audio signal semantic features are the same, obtaining the second non-difference audio time; obtaining a semantic difference audio time when the first reference audio signal semantic feature and the first monitoring audio signal semantic feature are different; The semantic feature extraction model construction step includes: A semantic vector mapping space is configured, which is used to map characters into unique multidimensional mathematical vectors; According to the semantic vector mapping space, a semantic feature extraction loss function is constructed: , wherein, characterizes a semantic feature extraction loss value, characterizes a regulation coefficient, 0 characterizes a kth dimension identification vector, characterizes a kth dimension extraction vector, H characterizes a mathematical vector dimension; With the model number of the sound source device as a constraint, audio signal recording data sets and audio semantic identification data sets of the same model device are collected; With the audio semantic identification data set as supervision and the audio signal recording data set as input, the semantic feature extraction model is trained based on the semantic vector mapping space and the semantic feature extraction loss function; The second reference audio signal label and the second monitoring audio signal label are analyzed for interval deviation to obtain an interval difference audio time, including: extracting the reference pitch and the reference frequency of the second reference audio signal label; extracting the monitoring pitch and the monitoring frequency of the second monitoring audio signal label; When the reference pitch and the monitoring pitch are different, or / and the reference frequency and the monitoring frequency are different, the interval difference audio time is obtained.
2. The method of claim 1, wherein, The interval difference audio time, the third reference audio signal label and the third monitoring audio signal label are added to the audio analysis result and sent to the user end, and further including: obtaining a physical feature difference audio time of audio physical feature comparison of the reference audio feature time sequence information and the monitoring audio feature time sequence information, wherein the physical feature difference audio time has a fourth reference audio signal label and a fourth monitoring audio signal label; obtaining a semantic difference audio time of semantic deviation analysis of the first reference audio signal label and the first monitoring audio signal label, wherein the semantic difference audio time has a fifth reference audio signal label and a fifth monitoring audio signal label; According to the interval difference audio time, the third reference audio signal label and the third monitoring audio signal label, an interval deviation ternary array is constructed, according to the physical feature difference audio time, the fourth reference audio signal label and the fourth monitoring audio signal label, a physical feature deviation ternary array is constructed, and according to the semantic difference audio time, the fifth reference audio signal label and the fifth monitoring audio signal label, a semantic deviation ternary array is constructed; The interval deviation ternary array, the physical feature deviation ternary array and the semantic deviation ternary array are added to the audio analysis result and sent to the user end.
3. The method of claim 1, wherein, The audio physical feature comparison function is: An audio physical feature comparison function is constructed: , , , , , , wherein, characterizing the reference audio feature timing information, characterizing the Q attribute reference audio feature timing information, characterizing the monitoring audio feature timing information, characterizing the Q attribute monitoring audio feature timing information, characterizing the i-th attribute reference audio feature timing information, characterizing the i-th attribute monitoring audio feature timing information, N characterizing the number of timing information of the i-th attribute reference audio feature timing information, M characterizing the number of timing information of the i-th attribute monitoring audio feature timing information, characterizing the adjustment parameter, > 1, characterizing the i-th attribute physical feature deviation, characterizing the audio physical feature global deviation coefficient. 4. An audio analysis system for differential localization, characterized by The system includes a user end, and the system and an audio sensor are in communication connection, and the audio sensor is detachably installed on a sound source device, including: A reference audio feature time sequence information obtaining module is configured to extract audio physical features in response to a first audio signal uploaded by the user end to obtain reference audio feature time sequence information; The monitoring audio feature timing information obtaining module is configured to receive a second audio signal collected by an audio sensor to extract audio physical features and obtain monitoring audio feature timing information; The audio physical feature comparison module is configured to compare the reference audio feature timing information and the monitoring audio feature timing information in terms of audio physical features to obtain a first non-difference audio time point, wherein the first non-difference audio time point has a first reference audio signal label and a first monitoring audio signal label; The semantic deviation analysis module is configured to analyze the first reference audio signal label and the first monitoring audio signal label in terms of semantics to obtain a second non-difference audio time point, wherein the second non-difference audio time point has a second reference audio signal label and a second monitoring audio signal label; The interval deviation analysis module is configured to analyze the second reference audio signal label and the second monitoring audio signal label in terms of intervals to obtain an interval difference audio time point, wherein the interval difference audio time point has a third reference audio signal label and a third monitoring audio signal label; The audio analysis result sending module is configured to add the interval difference audio time point, the third reference audio signal label and the third monitoring audio signal label into an audio analysis result and send the audio analysis result to a user terminal.
5. An electronic device, comprising: The memory is configured to store a computer software program; The processor is configured to read and execute the computer software program, thereby realizing the audio analysis method for difference positioning according to any one of claims 1 to 3. The storage medium stores a computer software program, and the computer software program is executed by the processor to realize the audio analysis method for difference positioning according to any one of claims 1 to 3.
6. A non-transitory computer-readable storage medium, comprising:
Citation Information
Patent Citations
Multimodal-based speech recording analysis method and system for people with language disorder
CN117831573A
Sound source positioning model training method, sound source object positioning method and related device
CN118675507A