Voice coding data processing method, device, equipment and chip
By pre-decoding the voice encoded data and adjusting the frame header information, spoofed frames are identified and corrected, thus solving the voice noise problem caused by network errors and improving the voice decoding quality and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNISOC CHONGQING TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-17
AI Technical Summary
During voice transmission, network transcoding or transmission errors may cause the frame content to differ from the frame header information, resulting in noise in the decoded voice data and affecting the user experience.
By pre-decoding the speech encoded data, spoofed frames are identified and the frame header information is adjusted to match the actual frame type, thus preventing the decoder from decoding the corrupted encoded content as noise.
It improves voice decoding quality, enhances the user experience, and ensures the clarity and continuity of the output voice.
Smart Images

Figure CN121884833A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, device and chip for processing voice encoded data. Background Technology
[0002] With the continuous development of communication technology, voice processing plays a crucial role in communication systems. During voice transmission, the transmitting end typically encodes the voice signal to generate voice-encoded data and sends it upstream. Then, after transcoding on the network side, the downstream receiving end decodes the voice-encoded data output from the network side using a decoder, thus realizing the effective transmission of audio data between different networks and devices.
[0003] However, in actual transmission, network transcoding or transmission errors can easily lead to differences between the acquired frame content and the original encoded frame, while the frame header information still records it as a normal frame. This causes the downlink decoder to decode it as a normal frame, resulting in noise in the decoded audio data and affecting the user experience. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, device, and chip for processing voice encoded data to address the aforementioned technical problems, thereby avoiding noise in the decoded voice data and improving the user experience.
[0005] Firstly, this application provides a method for processing speech encoded data, including:
[0006] The speech encoded data is pre-decoded, and the spoofed frame after pre-decoding is determined.
[0007] Adjust the frame header information of the corresponding spoofed frame in the speech coding data to the first frame header information. The first frame header information is used to indicate that the corresponding frame is abnormal.
[0008] The adjusted speech coding data is input to the decoder for speech decoding.
[0009] In one embodiment, pre-decoding the speech encoded data and determining the pre-decoded masquerading frame includes: pre-decoding the speech encoded data to extract target speech features corresponding to different target frames after pre-decoding; for each target frame, determining the target frame type of the target frame based on the target speech features of the target frame, wherein the target frame type includes a masquerading frame or a normal frame.
[0010] In one embodiment, the target speech features include spectral correlation features and / or energy features; accordingly, based on the target speech features of the target frame, the target frame type of the target frame is determined, including at least one of the following: in response to the frame energy change value in the energy features being greater than a preset energy change threshold, and the spectral correlation features satisfying a preset camouflage frame condition, the target frame type of the target frame is determined to be a camouflage frame; in response to the frame energy change value in the energy features being greater than a preset energy change threshold, the target frame type of the target frame is determined to be a camouflage frame; in response to the spectral correlation features satisfying the preset camouflage frame condition, the target frame type of the target frame is determined to be a camouflage frame.
[0011] In one embodiment, the preset energy change threshold is determined according to the following steps: obtaining the reference energy smoothing value corresponding to the previous frame of the target frame; determining the target energy smoothing value corresponding to the target frame based on the frame energy value and the reference energy smoothing value in the energy features; and determining the preset energy change threshold corresponding to the target frame based on the target energy smoothing value.
[0012] In one embodiment, a preset camouflage frame condition corresponds to at least one camouflage frame type; the preset camouflage frame condition includes a formant sub-condition and a filtered energy ratio sub-condition; the spectral correlation features include line spectrum frequencies and linear prediction coefficients; correspondingly, determining that the spectral correlation features satisfy the preset camouflage frame condition includes: determining the formant data corresponding to the target frame based on the line spectrum frequencies; and determining the frame energy ratio relationship of the target frame before and after filtering based on the linear prediction coefficients; for each camouflage frame type, in response to the formant data satisfying the formant sub-condition corresponding to the camouflage frame type and the frame energy ratio relationship satisfying the filtered energy ratio sub-condition corresponding to the camouflage frame type, determining that the spectral correlation features satisfy the preset camouflage frame condition.
[0013] In one embodiment, the target speech features include spectral correlation features and / or energy features; accordingly, based on the target speech features of the target frame, the target frame type of the target frame is determined, including at least one of the following: in response to the frame energy change value in the energy features not being greater than a preset energy change threshold, and the spectral correlation features not meeting a preset masquerading frame condition, the target frame type of the target frame is determined to be a normal frame; in response to the previous frame of the target frame being a normal frame, and the frame energy value in the energy features not being greater than a preset frame energy threshold, the target frame type of the target frame is determined to be a normal frame; in response to the previous frame of the target frame being a normal frame, and the frame energy change value in the energy features not being greater than a preset energy change threshold, the target frame type of the target frame is determined to be a normal frame; in response to the previous frame of the target frame being a normal frame, and the linear predicted reflection coefficient in the spectral correlation features not being greater than a preset coefficient threshold, the target frame type of the target frame is determined to be a normal frame.
[0014] In one embodiment, the method further includes: in response to the number of target frames that are consecutively spoofed frames before the target frame being greater than a preset frame number threshold, determining that the target frame type of the target frame is a normal frame.
[0015] Secondly, this application also provides a device for processing speech encoded data, comprising:
[0016] The first determining module is used to pre-decode the speech encoded data and determine the pre-decoded spoofed frame.
[0017] The adjustment module is used to adjust the frame header information of the corresponding spoofed frame in the speech coding data to the first frame header information, which is used to indicate the corresponding frame is abnormal.
[0018] The adjusted speech coding data is input to the decoder for speech decoding.
[0019] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program in accordance with the steps of the method provided in the first aspect above.
[0020] Fourthly, this application also provides a chip, including a processor and a communication interface, wherein the processor is configured to cause the chip to perform the steps of the method provided in the first aspect above.
[0021] Fifthly, this application also provides a chip module, including a communication module, a power module, a storage module, and a chip, wherein:
[0022] The power module is used to provide power to the chip module;
[0023] Storage modules are used to store data and instructions;
[0024] The communication module is used for internal communication within the chip module, or for communication between the chip module and external devices;
[0025] The chip is used to perform the steps of the method provided in the first aspect above.
[0026] Sixthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect above.
[0027] In a seventh aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method provided in the first aspect above.
[0028] The aforementioned method, apparatus, device, and chip for processing speech encoded data pre-decode the speech encoded data, thereby pre-decoding the speech encoded data in the coding domain into analyzable speech features, providing a data foundation for the subsequent determination of spoofed frames. By determining the pre-decoded spoofed frames, distorted spoofed frames generated due to network transmission or network transcoding errors can be identified. By adjusting the frame header information of the corresponding spoofed frame in the speech encoded data to the first frame header information, which indicates the corresponding frame anomaly, the frame header information in the speech encoded data is updated, ensuring that the frame header information matches the actual frame type of the frame data. By inputting the adjusted speech encoded data into the decoder, the decoder can perform speech decoding based on the actual and correct frame type, avoiding directly decoding corrupted encoded content into noise, improving the output speech quality, and enhancing the user experience. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is an application environment diagram of a speech coding data processing method in one embodiment;
[0031] Figure 2A This is a flowchart illustrating a method for processing speech encoded data in one embodiment;
[0032] Figure 2B This is a schematic diagram of a single formant camouflage frame in one embodiment;
[0033] Figure 2C This is a schematic diagram of a white noise camouflage frame in one embodiment;
[0034] Figure 2D This is a schematic diagram of a multi-formant camouflage frame in one embodiment;
[0035] Figure 2E This is a schematic diagram of a normal frame in one embodiment;
[0036] Figure 2F This is a schematic diagram of the processing chain of a speech coding data processing method in one embodiment;
[0037] Figure 3A This is a timing diagram of a speech coding data processing method in one embodiment;
[0038] Figure 3B This is a schematic diagram of the detection process for spoofed frames in one embodiment;
[0039] Figure 4 This is a speech encoding data processing device in one embodiment;
[0040] Figure 5 This is an internal structural diagram of a computer device in one embodiment;
[0041] Figure 6 This is an internal structure diagram of a chip module in one embodiment. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0043] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various features, but these features are not limited by these terms. These terms are only used to distinguish the first feature from the second feature. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0044] The speech coding data processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, the sender 101, server 103, and receiver 102 communicate via a network.
[0045] For example, the sending end 101 can be, but is not limited to, various personal computers, laptops, smartphones, or tablets. The receiving end 102 can be, but is not limited to, various personal computers, laptops, smartphones, or tablets. The server 103 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0046] The transmitting end 101 is mainly used to encode the input voice signal to obtain voice encoded data and send it to the server 103; the server is mainly used to transmit the voice encoded data to the receiving end 102; the receiving end 102 is mainly used to decode the voice encoded data to obtain the decoded voice signal. For example, the server can also perform network transcoding on the voice encoded data to adapt to different encoding formats or network conditions.
[0047] Understandably, during actual transmission, network transcoding or transmission errors can easily cause differences between the frame content obtained by the receiving end 102 and the original encoded frame, but the frame header information still records it as a normal frame. This leads the subsequent decoder to decode it as a normal frame, resulting in noise in the decoded audio data and affecting the user experience. Based on this, this application provides a method for processing audio encoded data. By pre-decoding the audio encoded data, the audio encoded data in the encoding domain can be pre-decoded into analyzable audio features, providing a data basis for the subsequent determination of spoofed frames. By determining the pre-decoded spoofed frames, spoofed frames with distorted content generated due to network transmission or transcoding errors can be identified. By adjusting the frame header information of the corresponding spoofed frame in the audio encoded data to the first frame header information, which indicates the corresponding frame is abnormal, the frame header information in the audio encoded data is updated, ensuring that the frame header information matches the actual frame type of the frame data. By inputting the adjusted speech encoding data into the decoder, the decoder can perform speech decoding based on the actual and correct frame type, avoiding the direct decoding of corrupted encoded content into noise, thus improving the output speech quality and enhancing the user experience.
[0048] In one exemplary embodiment, such as Figure 2A As shown, a method for processing speech encoded data is provided, which can be applied to... Figure 1 Taking the receiver 102 in the example, or a chip / chip module with data processing capabilities, as an example, the following steps are included:
[0049] S210. Perform pre-decoding processing on the speech encoded data and determine the spoofed frame after pre-decoding.
[0050] Here, voice-encoded data can be understood as the acquired voice-encoded data after the voice signal has been encoded by the transmitting end. For example, voice-encoded data can be voice-encoded data after network transmission and / or transcoding processing.
[0051] For example, the initial frame header information corresponding to the speech encoded data indicates that each target frame corresponding to the speech encoded data is a normal frame. Understandably, during network transmission or transcoding, the target frame may become abnormal due to network transmission errors or transcoding errors, resulting in the actual frame type of the target frame (i.e., the abnormal frame type) not matching the initial frame header information, causing subsequent speech decoding to decode as noise.
[0052] In this context, spoofed frames, also known as abnormal frames, are used to represent frame data that is distorted, abrupt, or corrupted, failing to conform to the natural characteristics of speech. Normal frames, on the other hand, represent frame data that conforms to the natural characteristics of speech. Referring to the foregoing, the initial frame header information corresponding to the acquired speech-coded data indicates that each target frame corresponding to the speech-coded data is a normal frame. Therefore, spoofed frames can also be understood as frame data whose frame type does not match the corresponding frame header information and that is distorted, abrupt, or corrupted, failing to conform to the natural characteristics of speech.
[0053] Pre-decoding can be understood as a pre-decoding operation performed to analyze frame features before the speech encoded data is decoded and played back.
[0054] For example, target speech features for different target frames can be decoded from the speech coding data according to the coding format of the speech coding data. Optionally, the speech coding data can be input into a pre-decoding module to obtain target speech features for different target frames.
[0055] Understandably, by pre-decoding the speech encoded data, the target speech features of different target frames of the target speech data corresponding to the speech encoded data can be obtained. Compared to first decoding the speech encoded data to obtain the target speech data, and then extracting the target speech features from the obtained target speech data, this method can improve the efficiency of target speech feature acquisition and reduce redundant data processing. Of course, in other embodiments, the target speech data can also be obtained by first decoding the speech encoded data, and then the target speech features can be extracted from the target speech data; this application does not impose any limitations on this method.
[0056] In an optional embodiment, the target speech features can be input into a spoofing frame recognition model to obtain spoofing frames in the target speech data. The spoofing frame recognition model can be a traditional machine learning model or a neural network model, and this application does not impose any limitations on it.
[0057] In another optional embodiment, the speech encoded data can be pre-decoded to extract the target speech features corresponding to different target frames after pre-decoding; for each target frame, the target frame type is determined according to the target speech features of the target frame, and the target frame type includes a bad frame or a good frame.
[0058] Optionally, the target speech features may include spectral correlation features and / or energy features. For example, spectral correlation features may include at least one of line spectral frequencies (LSF), linear prediction coefficients (LPC), and linear prediction reflection coefficients; energy features may include at least one of frame energy values and frame energy variation values. Here, line spectral frequencies (LSF) are used to characterize the spectral distribution; linear prediction coefficients (LPC) can be used as filter coefficients, i.e., LPC filters based on linear prediction coefficients (LPC) can be used to filter the speech-coded data; linear prediction coefficients (LPC) and linear prediction reflection coefficients can be obtained by conversion based on line spectral frequencies (LSF), and the conversion method can employ traditional mathematical transformations, which are not limited in this application.
[0059] Optionally, the spectral correlation features may also include line spectral pairs (LSPs). The target speech features may also include at least one of pitch period, gene gain, and codebook gain. This application does not impose any limitations on the specific feature types of the target speech features.
[0060] In an optional embodiment, the camouflage frame type may include at least one of a single formant camouflage frame, a multi-formant camouflage frame, and a white noise-like camouflage frame.
[0061] To facilitate understanding, the above three types of camouflage frames are illustrated below with reference to the accompanying drawings.
[0062] refer to Figure 2B The diagram shows a single-resonant camouflage frame. Wherein, Figure 2B The image shows the time-domain and frequency-domain characteristics of a single-resonant camouflage frame. A single-resonant camouflage frame refers to a frame where energy is concentrated on a single resonant peak with a relatively low frequency. From... Figure 2B The data shows that the formant appears at 100Hz with an amplitude of 40dB. Normal human speech does not exhibit a single formant at extremely low frequencies and high energy; the subjective listening experience of a single formant masquerade frame is that of an electronic sound.
[0063] refer to Figure 2C The image shows a schematic diagram of a white noise camouflage frame. Among them, Figure 2C The diagram illustrates the time-domain and frequency-domain characteristics of a white noise-like camouflage frame. The white noise-like camouflage frame exhibits a uniform frequency-domain energy distribution without obvious concentrated regions, and its spectral correlation characteristics are less regular than those of white noise; hence, it is called a white noise-like camouflage frame.
[0064] refer to Figure 2D The diagram shows a multi-resonant camouflage frame. Among them, Figure 2D The image shows the time-domain and frequency-domain characteristics of a multi-formant camouflage frame. The characteristics of a multi-formant camouflage frame fall between those of a single-formant camouflage frame and a white noise-like camouflage frame; it typically has multiple formants with high energy.
[0065] refer to Figure 2E The image shown is a schematic diagram of a normal frame. Among them, Figure 2E The image shows the vowel temporal domain signal and vowel audio domain characteristics of a normal frame. From... Figure 2E As can be seen from the figure, the time-domain signal of the normal frame has obvious periodicity, and the fundamental frequency of the frequency-domain signal is obvious. As shown in the figure, the resonance peak appears at around 300Hz, and there are multiple resonance peaks with small energy differences among them.
[0066] Among these, spoofed frames exhibit a degree of randomness, lacking a single, specific time-domain or frequency-domain feature. Using a single time-domain or frequency-domain feature for spoofed frame detection results in high false positive and false negative rates. After analyzing hundreds of spoofed frame speech data, the spoofed frames were categorized into at least three main types: single-formant spoofed frames, multi-formant spoofed frames, and white noise-like spoofed frames. These three types of spoofed frames each have their own characteristics, but their common feature is a sudden increase in frame energy.
[0067] Table 1 Summary of characteristics of different camouflage frames
[0068]
[0069] Table 1 summarizes the characteristics of different camouflage frames. Table 1 provides examples of the periodicity, formant characteristics, spectral energy characteristics, and auditory characteristics corresponding to different camouflage frame types.
[0070] In one optional embodiment, preset camouflage frame conditions can be constructed for each camouflage frame type; the target frame type is determined by determining whether the spectral correlation characteristics of the target frame satisfy different preset camouflage frame conditions. Furthermore, the target frame type can be determined by combining energy characteristics and / or the frame types of historical frames corresponding to the target frame, thereby improving the accuracy of determination. Of course, in some implementations, the target frame type can also be determined by comparing energy characteristics with corresponding thresholds or based on the frame types of historical frames.
[0071] S220. Adjust the frame header information of the corresponding spoofed frame in the speech coding data to the first frame header information, which is used to indicate the corresponding frame is abnormal; wherein, the adjusted speech coding data is used to input to the decoder to perform speech decoding on the adjusted speech coding data.
[0072] In an optional embodiment, for a spoofed frame in the speech coding data, the decoder can obtain the previous frame of the spoofed frame and perform error concealment processing on the spoofed frame, i.e. packet loss concealment processing, so as to maintain the auditory continuity of the decoded target speech and avoid the generation of noise.
[0073] refer to Figure 2F The diagram illustrates the processing chain of a voice encoded data processing method. It involves acquiring voice encoded data transmitted from a modem, performing spoofing frame detection and adjustment on the voice encoded data, and then using a decoder or audio digital signal processor (ADSP) to perform voice decoding on the adjusted voice encoded data.
[0074] For example, the spoofing frame detection and adjustment process may include: pre-decoding the speech encoded data; calculating speech features, i.e., target speech features corresponding to different target frames; detecting spoofing frames; and modifying the frame header of the speech encoded data, i.e., adjusting the frame header information of the corresponding spoofing frame in the speech encoded data.
[0075] For example, the adjusted speech coding data can be filtered during the decoding process. Specifically, the speech excitation can be synthesized into speech based on a synthesis filter.
[0076] The aforementioned method for processing speech encoded data pre-decodes the data, transforming it into analyzable speech features and providing a data foundation for identifying subsequent spoofed frames. By identifying the pre-decoded spoofed frames, it becomes possible to identify distorted frames generated due to network transmission or transcoding errors. By adjusting the frame header information of the corresponding spoofed frame in the speech encoded data to the first frame header information (indicating frame anomalies), the frame header information in the speech encoded data is updated, ensuring it matches the actual frame type. By inputting the adjusted speech encoded data into the decoder, the decoder can perform speech decoding based on the actual and correct frame type, avoiding the direct decoding of corrupted encoded content into noise, thus improving output speech quality and enhancing the user experience.
[0077] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment in which the steps for determining the target frame type are refined.
[0078] In an optional embodiment, the target speech features include spectral correlation features and / or energy features; referring to the foregoing, the target frame type can be determined based on the target speech features of the target frame. For ease of understanding, the steps for determining the target frame type are illustrated below using methods a1, a2, and a3 as examples.
[0079] Method a1: If the frame energy change value in the energy feature is greater than the preset energy change threshold and the spectrum correlation feature meets the preset camouflage frame condition, the target frame type of the target frame is determined to be a camouflage frame.
[0080] Method a2: When the frame energy change value in the energy feature is greater than the preset energy change threshold, the target frame type of the target frame is determined to be a camouflaged frame.
[0081] Method a3: In response to the spectral correlation features satisfying the preset camouflage frame conditions, the target frame type of the target frame is determined to be a camouflage frame.
[0082] Among them, the frame energy change value is used to characterize the degree of frame energy change between the target frame and the previous frame.
[0083] The preset camouflage frame conditions can be understood as the judgment conditions pre-set for identifying camouflage frames based on the spectrum-related feature dimensions. Referring to the foregoing, the preset camouflage frame conditions can correspond to different camouflage frame types, which may include at least one of single-formant camouflage frames, multi-formant camouflage frames, and white noise-like camouflage frames.
[0084] The following is an exemplary description of the steps for determining the preset energy change threshold. It should be noted that this should not be construed as a limitation on the specific process for determining the preset energy change threshold. For example, the preset energy change threshold corresponding to the target frame can be determined based on the target energy smoothing value corresponding to the target frame.
[0085] In an optional embodiment, a reference energy smoothing value corresponding to the previous frame of the target frame can be obtained; a target energy smoothing value corresponding to the target frame can be determined based on the frame energy value and the reference energy smoothing value in the energy features; and a preset energy change threshold corresponding to the target frame can be determined based on the target energy smoothing value. The preset energy change threshold can be determined based on a preset matching relationship and the target energy smoothing value. Optionally, the larger the target energy smoothing value, the larger the preset energy change threshold.
[0086] Optionally, the target energy smoothing value corresponding to the nth target frame can be determined according to the following formula:
[0087]
[0088] in, This represents the energy smoothing value corresponding to the nth target frame, i.e., the target energy smoothing value; Indicates the smoothing coefficient; The energy smoothing value corresponding to the (n-1)th target frame, i.e., the reference energy smoothing value; This represents the frame energy value corresponding to the nth target frame.
[0089] Optionally, the smoothing coefficient can be determined based on the frame energy value corresponding to the nth target frame and the energy smoothing value corresponding to the (n-1)th target frame. .
[0090] For example, the smoothing coefficient can be determined according to the following formula. :
[0091]
[0092] in, This represents the first preset smoothing coefficient; This represents the second preset smoothing coefficient; > .
[0093] The following is an exemplary description of the matching process for preset camouflage frame conditions. It should be noted that this should not be construed as a limitation on the specific matching process.
[0094] In an optional embodiment, a preset camouflage frame condition corresponds to at least one camouflage frame type; the preset camouflage frame condition includes a formant sub-condition and a filtered energy ratio sub-condition; the spectral correlation features include line spectrum frequencies and linear prediction coefficients; correspondingly, the formant data corresponding to the target frame can be determined based on the line spectrum frequencies; and the frame energy ratio relationship of the target frame before and after filtering based on the linear prediction coefficients can be determined; for each camouflage frame type, in response to the formant data satisfying the formant sub-condition corresponding to the camouflage frame type and the frame energy ratio relationship satisfying the filtered energy ratio sub-condition corresponding to the camouflage frame type, it is determined that the spectral correlation features satisfy the preset camouflage frame condition.
[0095] The resonance peak data may include at least one of the following: the number of resonance peaks, the energy distribution of resonance peaks, the shape of resonance peaks, and the LSF (line spectral frequency) distance representing the intensity of resonance peaks.
[0096] The frame energy ratio relationship may include the target ratio R; the target ratio R is used to characterize the ratio of the frame energy of the target frame after filtering based on linear prediction coefficients to that before filtering.
[0097] Table 2 Summary of conditions corresponding to single formant camouflage frames
[0098]
[0099] Table 2 shows a summary of the conditions corresponding to a single-resonant-peak camouflage frame. Table 2 lists the filter energy ratio sub-condition A1 and the resonant-peak sub-condition B1 for the single-resonant-peak camouflage frame. The preset camouflage frame conditions for a single-resonant-peak camouflage frame may also include the reflection coefficient sub-condition C1; the linear predicted reflection coefficient in the spectral correlation features may include multi-order reflection coefficients, including the first-order reflection coefficient K0.
[0100] Where R represents the target ratio; the first ratio threshold R1, the preset distance threshold LD, and the first reflection coefficient threshold K vowel The threshold value can be set by technicians based on their needs or experience, or determined through extensive experimentation; this application does not impose any limitations in this regard. For example, the first reflectance coefficient threshold K... vowel It can be -0.9.
[0101] For example, if the frame energy change value is greater than a preset energy change threshold and the filter energy ratio sub-condition A1 and the resonance peak sub-condition B1 are satisfied, the target frame can be determined to be a single resonance peak masquerading frame.
[0102] For example, if the frame energy change value is greater than a preset energy change threshold and the filter energy ratio sub-condition A1, the resonant peak sub-condition B1, and the reflection coefficient sub-condition C1 are satisfied, the target frame can be determined as a single resonant peak camouflage frame.
[0103] Table 3 Summary of conditions corresponding to white noise camouflage frames
[0104]
[0105] Table 3 shows a summary of the conditions corresponding to the white noise camouflage frame. Table 3 lists the filter energy ratio sub-condition A2 and the resonance peak sub-condition B2 for the white noise camouflage frame. Furthermore, the preset camouflage frame conditions for the white noise camouflage frame may also include the reflection coefficient condition C2. The linear predicted reflection coefficient in the spectral correlation features may include multi-order reflection coefficients, including the first-order reflection coefficient K0.
[0106] Where R represents the target ratio; the second proportional threshold R2 and the second reflection coefficient threshold K white The threshold value can be set by technicians according to their needs or experience, or determined through extensive experiments; this application does not impose any limitations on this. R2 < R1. For example, the second reflection coefficient threshold K... white It can be -0.5.
[0107] For example, if the frame energy change value is greater than a preset energy change threshold and the filter energy ratio sub-condition A2 and the resonance peak sub-condition B2 are satisfied, the target frame can be determined as a white noise camouflage frame.
[0108] For example, a target frame can be determined to be a white noise camouflage frame if the frame energy change value is greater than a preset energy change threshold and the filter energy ratio sub-condition A2, the resonance peak sub-condition B2, and the reflection coefficient condition C2 are satisfied.
[0109] Table 4 Summary of conditions corresponding to multi-resonant camouflage frames
[0110]
[0111] Table 4 shows a summary of the conditions corresponding to multi-resonant camouflage frames. Specifically, Table 4 lists the filter energy ratio sub-condition A3 and the resonant peak sub-condition B3 for multi-resonant camouflage frames. Furthermore, the preset camouflage frame conditions for multi-resonant camouflage frames may also include the reflection coefficient condition C3.
[0112] For example, if the frame energy change value is greater than the preset energy change threshold and the filter energy ratio sub-condition A3 and the resonance peak sub-condition B3 are satisfied, the target frame can be determined to be a multi-resonant peak camouflage frame.
[0113] For example, if the frame energy change value is greater than a preset energy change threshold and the filter energy ratio sub-condition A3, the resonant peak sub-condition B3, and the reflection coefficient sub-condition C3 are satisfied, the target frame can be determined as a multi-resonant peak camouflage frame.
[0114] For example, the preset frequency threshold can be 1.25kHz; the high-frequency resonant peak can be a resonant peak with a frequency greater than 1.25kHz, and the low-frequency resonant peak can be a resonant peak with a frequency not greater than 1.25kHz.
[0115] In an optional embodiment, the target frame type of the target frame can be determined to be a normal frame in response to the fact that the frame energy change value in the energy features is not greater than a preset energy change threshold and the spectral correlation features do not meet the preset camouflage frame condition. The preset energy change threshold and the preset camouflage frame condition have been described above and will not be repeated here.
[0116] Based on the above embodiments, the target frame type can also be determined by combining the frame types of the historical frames corresponding to the target frame. For example, the target frame type can be determined by a masquerading frame detection state machine.
[0117] In an optional embodiment, the target frame type of the target frame can be determined to be a normal frame in response to the previous frame of the target frame being a normal frame and the frame energy value in the energy feature not being greater than a preset frame energy threshold.
[0118] In an optional embodiment, the target frame type of the target frame can be determined to be a normal frame in response to the previous frame of the target frame being a normal frame and the frame energy change value in the energy feature not being greater than a preset energy change threshold.
[0119] In an optional embodiment, the target frame type of the target frame can be determined to be a normal frame in response to the previous frame of the target frame being a normal frame and the linear predicted reflection coefficient in the spectrum correlation feature not being greater than a preset coefficient threshold.
[0120] In an optional embodiment, the target frame type of the target frame can be determined to be a normal frame in response to the number of target frames that are consecutively spoofed frames before the target frame being greater than a preset frame number threshold.
[0121] Optionally, in response to the previous frame of the target frame being a spoofed frame, the number of consecutive target frames that were spoofed frames before the target frame can be obtained; in response to the number of target frames being greater than a preset frame number threshold, the target frame type of the target frame can be determined to be a normal frame. For example, the distribution of historical spoofed frames can be obtained, and the number of historical frames in which spoofed frames appeared consecutively can be obtained; the maximum number of frames N in the historical frame count can be used as the preset frame number threshold.
[0122] In an optional embodiment, the target frame type of the target frame can be determined to be a normal frame in response to the fact that the previous frame of the target frame is a spoofed frame, the frame energy change value in the energy feature is not greater than a preset energy change threshold, and the spectrum correlation feature does not meet the preset spoofed frame condition.
[0123] In an optional embodiment, if the previous frame of the target frame is a normal frame or a camouflaged frame, and the frame energy change value in the energy feature is greater than a preset energy change threshold, and the spectrum correlation feature meets the preset camouflaged frame condition, then the target frame type of the target frame is determined to be a camouflaged frame.
[0124] Based on the above embodiments, in some embodiments, the method for processing voice encoded data can be implemented through interaction between the sending end, the server, and the receiving end.
[0125] refer to Figure 3A The diagram shown is a timing diagram of a method for processing speech encoded data, including the following steps:
[0126] S301. The transmitting end encodes the voice signal to generate voice encoded data.
[0127] S302, The sending end sends voice-encoded data to the server.
[0128] S303, The server performs network transcoding on the voice encoded data.
[0129] S304. The server sends the transcoded voice data to the receiving end.
[0130] S305. The receiving end performs pre-decoding processing on the speech coding data and extracts the target speech features corresponding to different target frames after pre-decoding processing; the target speech features include spectrum correlation features and / or energy features; the spectrum correlation features include line spectrum frequency and linear prediction coefficient.
[0131] Among these, the spectral correlation features may also include the reflection coefficient.
[0132] S306. For each target frame, the receiver determines the resonance peak data corresponding to the target frame based on the line spectrum frequency when the frame energy change value in the energy characteristics of the target frame is greater than the preset energy change threshold, and determines the frame energy ratio relationship of the target frame before and after filtering based on linear prediction coefficients.
[0133] S307. For each spoofing frame type, the receiving end determines the corresponding target frame as a spoofing frame in response to the formant data satisfying the formant sub-condition corresponding to the spoofing frame type and the frame energy ratio satisfying the filter energy ratio sub-condition corresponding to the spoofing frame type.
[0134] S308. The receiving end adjusts the frame header information of the corresponding spoofed frame in the voice coding data to the first frame header information. The first frame header information is used to indicate that the corresponding frame is abnormal.
[0135] S309. The receiving end performs voice decoding on the adjusted voice encoded data through the decoder.
[0136] refer to Figure 3B The diagram illustrates the detection process for spoofed frames. It determines whether the number of consecutive spoofed frames exceeds a threshold, i.e., a preset frame count threshold. If so, the current frame is determined to be a normal frame. If not, the detection threshold and the speech features of the current frame are calculated based on the previous frame's state, and a two-level detection process is initiated. The first level detects abrupt changes in energy between adjacent frames, while the second level checks whether the current frame meets the preset spoofing frame conditions. Specifically:
[0137] Level 1 detection: Determine whether the frame energy changes abruptly based on speech features, i.e., whether the frame energy change value is greater than the preset energy change threshold; if not, determine that the current frame is a normal frame; otherwise, proceed to Level 2 detection.
[0138] Second-level detection: Determine whether the speech features meet the preset camouflage frame conditions corresponding to single formant camouflage frames, multi-formant camouflage frames, or white noise-like camouflage frames; if at least one preset camouflage frame condition is met, the current frame is determined to be a camouflage frame; if none of the preset camouflage frame conditions are met, the current frame is determined to be a normal frame.
[0139] The first spoofed frame shows a significant change in frame energy compared to the previous normal frame, allowing it to pass the first level of detection. However, subsequent spoofed frames show no significant change in frame energy compared to the previous one, potentially leading to missed detections. Therefore, after detecting the first spoofed frame, the subsequent detection thresholds are adjusted accordingly. For example, based on the frame energy value and a reference energy smoothing value, a target energy smoothing value is determined for the target frame; and based on the target energy smoothing value, a preset energy change threshold is determined for the target frame.
[0140] For AMR encoding, each frame is 20ms long. In the tests conducted so far, the longest spoofed frame has been 120ms. Therefore, a consecutive spoofed frame threshold of 7 can be set. After detecting 7 consecutive spoofed frames, the next frame will not be detected as a spoofed frame, reducing false detections and speech distortion.
[0141] In the above process, by performing two-level detection on speech features, detection efficiency and accuracy can be improved, while false positive and false negative rates can be reduced. By calculating the detection threshold based on the state of the previous frame, the preset energy change threshold can be adaptively determined according to the actual situation, thereby further improving detection accuracy. By setting a maximum limit on the number of spoofed frames, false positives and speech distortion can be reduced.
[0142] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0143] Based on the same inventive concept, this application also provides a speech encoded data processing apparatus for implementing the speech encoded data processing method described above. This apparatus can be applied to or integrated into a chip or chip module, for example. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more speech encoded data processing apparatus embodiments provided below can be found in the limitations of the speech encoded data processing method described above, and will not be repeated here.
[0144] In one exemplary embodiment, such as Figure 4 As shown, a speech encoding data processing apparatus is provided, comprising: a first determining module 410 and an adjusting module 420, wherein:
[0145] The first determining module 410 is used to pre-decode the speech encoded data and determine the pre-decoded spoofed frame.
[0146] The adjustment module 420 is used to adjust the frame header information of the corresponding spoofed frame in the speech encoding data to the first frame header information, which is used to indicate the corresponding frame is abnormal; wherein, the adjusted speech encoding data is used to input to the decoder to perform speech decoding on the adjusted speech encoding data.
[0147] In one embodiment, the first determining module 410 includes: an extraction unit, configured to pre-decode the speech encoded data to extract target speech features corresponding to different target frames after pre-decoding; and a first determining unit, configured to determine the target frame type of each target frame based on the target speech features of the target frame, wherein the target frame type includes a masquerading frame or a normal frame.
[0148] In one embodiment, the target speech features include spectral correlation features and / or energy features; correspondingly, the first determining unit includes at least one of the following: a first determining subunit, configured to determine the target frame type of the target frame as a camouflaged frame in response to the frame energy change value in the energy features being greater than a preset energy change threshold and the spectral correlation features satisfying a preset camouflaged frame condition; a second determining subunit, configured to determine the target frame type of the target frame as a camouflaged frame in response to the frame energy change value in the energy features being greater than the preset energy change threshold; and a third determining subunit, configured to determine the target frame type of the target frame as a camouflaged frame in response to the spectral correlation features satisfying the preset camouflaged frame condition.
[0149] In one embodiment, the first determining unit further includes: an acquisition subunit, configured to acquire a reference energy smoothing value corresponding to the previous frame of the target frame; a fourth determining subunit, configured to determine a target energy smoothing value corresponding to the target frame based on the frame energy value and the reference energy smoothing value in the energy features; and a fifth determining subunit, configured to determine a preset energy change threshold corresponding to the target frame based on the target energy smoothing value.
[0150] In one embodiment, the preset camouflage frame condition corresponds to at least one camouflage frame type; the preset camouflage frame condition includes a formant sub-condition and a filtered energy ratio sub-condition; the spectral correlation features include line spectrum frequency and linear prediction coefficient; correspondingly, the first determining unit further includes: a sixth determining sub-unit, used to determine the formant data corresponding to the target frame based on the line spectrum frequency; a seventh determining sub-unit, used to determine the frame energy ratio relationship of the target frame before and after filtering based on the linear prediction coefficient; and an eighth determining sub-unit, used to determine, for each camouflage frame type, in response to the formant data satisfying the formant sub-condition corresponding to the camouflage frame type and the frame energy ratio relationship satisfying the filtered energy ratio sub-condition corresponding to the camouflage frame type, that the spectral correlation features satisfy the preset camouflage frame condition.
[0151] In one embodiment, the target speech features include spectral correlation features and / or energy features; correspondingly, the first determining unit includes at least one of the following: a ninth determining subunit, configured to determine the target frame type of the target frame as a normal frame in response to the frame energy change value in the energy features not being greater than a preset energy change threshold and the spectral correlation features not meeting a preset masquerading frame condition; a tenth determining subunit, configured to determine the target frame type of the target frame as a normal frame in response to the previous frame of the target frame being a normal frame and the frame energy value in the energy features not being greater than a preset frame energy threshold; an eleventh determining subunit, configured to determine the target frame type of the target frame as a normal frame in response to the previous frame of the target frame being a normal frame and the frame energy change value in the energy features not being greater than a preset energy change threshold; and a twelfth determining subunit, configured to determine the target frame type of the target frame as a normal frame in response to the previous frame of the target frame being a normal frame and the linear predicted reflection coefficient in the spectral correlation features not being greater than a preset coefficient threshold.
[0152] In one embodiment, the above apparatus further includes: a second determining device, configured to determine that the target frame type of the target frame is a normal frame in response to the number of target frames that are consecutively spoofed frames before the target frame being greater than a preset frame number threshold.
[0153] Regarding the modules / units included in the various devices and products described in the above embodiments, they can be software modules / units, hardware modules / units, or a combination of both. For example, for various devices and products applied to or integrated into a chip, all of their modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs that run on a processor integrated within the chip, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits; for various devices and products applied to or integrated into a chip module, all of their modules / units can be implemented using hardware methods such as circuits, and different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using hardware methods such as circuits. The components can be implemented using software programs that run on the processor integrated within the chip module. The remaining (if any) modules / units can be implemented using hardware methods such as circuits. For various devices and products applied to or integrated into the terminal, each of its components / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or in different components within the terminal. Alternatively, at least some modules / units can be implemented using software programs that run on the processor integrated within the terminal, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits.
[0154] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for processing voice-encoded data. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0155] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0156] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0157] Based on the same inventive concept, this application also provides a chip, including a processor coupled to a memory, for executing a computer program or instructions stored in the memory, and implementing the steps in the above method embodiments when the processor executes the computer program or instructions.
[0158] It is understood that the chip involved in the embodiments of this application may be a field-programmable gate array (FPGA), may include an application-specific integrated circuit (ASIC), may be a system on chip (SoC), may be a central processor unit (CPU), may be a network processor (NP), may be a digital signal processor (DSP), may be a microcontroller unit (MCU), may be a programmable logic device (PLD), or other integrated chips, etc.
[0159] Based on the same inventive concept, this application also provides a chip module, such as... Figure 6 As shown, the chip module includes a communication module, a power module, a storage module, and a chip. Among them:
[0160] The power module provides power to the chip module; the storage module stores data and instructions; the communication module enables internal communication within the chip module or communication between the chip module and external devices; this chip corresponds to the chip in the above chip embodiment. The implementation of this chip module can be found in the relevant content of the above chip embodiment, and will not be repeated here.
[0161] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0162] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0163] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0164] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0165] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for processing speech encoded data, characterized in that, The method includes: The speech encoded data is pre-decoded, and the spoofed frame after pre-decoding is determined. The frame header information corresponding to the spoofed frame in the speech encoding data is adjusted to the first frame header information, which is used to indicate that the corresponding frame is abnormal. The adjusted speech encoding data is input to the decoder to perform speech decoding.
2. The method according to claim 1, characterized in that, The step of pre-decoding the speech encoded data and determining the pre-decoded spoofed frame includes: The speech encoded data is pre-decoded to extract the target speech features corresponding to different target frames after pre-decoding; For each target frame, the target frame type is determined based on the target speech features of the target frame. The target frame type includes a masquerading frame or a normal frame.
3. The method according to claim 2, characterized in that, The target speech features include spectral correlation features and / or energy features; correspondingly, determining the target frame type of the target frame based on the target speech features of the target frame includes at least one of the following: If the frame energy change value in the energy feature is greater than a preset energy change threshold, and the spectrum correlation feature satisfies a preset camouflage frame condition, the target frame type of the target frame is determined to be a camouflage frame. If the frame energy change value in the energy feature is greater than a preset energy change threshold, the target frame type of the target frame is determined to be a camouflaged frame. In response to the spectral correlation features satisfying the preset camouflage frame conditions, the target frame type of the target frame is determined to be a camouflage frame.
4. The method according to claim 3, characterized in that, The preset energy change threshold is determined according to the following steps: Obtain the reference energy smoothing value corresponding to the previous frame of the target frame; Based on the frame energy value in the energy features and the reference energy smoothing value, determine the target energy smoothing value corresponding to the target frame; Based on the target energy smoothing value, a preset energy change threshold corresponding to the target frame is determined.
5. The method according to claim 3, characterized in that, The preset camouflage frame condition corresponds to at least one camouflage frame type; the preset camouflage frame condition includes a resonant peak sub-condition and a filter energy ratio sub-condition; the spectral correlation features include line spectrum frequencies and linear prediction coefficients; correspondingly, determining that the spectral correlation features satisfy the preset camouflage frame condition includes: Based on the line spectrum frequencies, determine the formant data corresponding to the target frame; and, Determine the frame energy ratio of the target frame before and after filtering based on the linear prediction coefficients; For each camouflage frame type, in response to the formant data satisfying the formant sub-condition corresponding to the camouflage frame type, and the frame energy ratio relationship satisfying the filter energy ratio sub-condition corresponding to the camouflage frame type, it is determined that the spectrum correlation feature satisfies the preset camouflage frame condition.
6. The method according to claim 2, characterized in that, The target speech features include spectral correlation features and / or energy features; correspondingly, determining the target frame type of the target frame based on the target speech features of the target frame includes at least one of the following: If the frame energy change value in the energy feature is not greater than a preset energy change threshold, and the spectrum correlation feature does not meet the preset camouflage frame condition, the target frame type of the target frame is determined to be a normal frame. In response to the previous frame of the target frame being a normal frame and the frame energy value in the energy feature not being greater than a preset frame energy threshold, the target frame type of the target frame is determined to be a normal frame. In response to the previous frame of the target frame being a normal frame and the frame energy change value in the energy feature not being greater than a preset energy change threshold, the target frame type of the target frame is determined to be a normal frame. In response to the previous frame of the target frame being a normal frame, and the linear predicted reflection coefficient in the spectrum correlation feature not being greater than a preset coefficient threshold, the target frame type of the target frame is determined to be a normal frame.
7. The method according to claim 6, characterized in that, The method further includes: In response to the number of target frames that are consecutively spoofed frames before the target frame being greater than a preset frame number threshold, the target frame type of the target frame is determined to be a normal frame.
8. A device for processing voice encoded data, characterized in that, The device includes: The first determining module is used to pre-decode the speech encoded data and determine the pre-decoded spoofed frame. The adjustment module is used to adjust the frame header information corresponding to the spoofed frame in the speech encoding data to the first frame header information, the first frame header information being used to indicate the corresponding frame is abnormal; The adjusted speech encoding data is input to the decoder to perform speech decoding.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A chip, characterized in that, The device includes a processor and a communication interface, the processor being configured to cause the chip to perform the steps of the method described in any one of claims 1 to 7.