Noise processing method and device, electronic equipment and computer readable storage medium

By acquiring the noise features and semantic features of audio and video streams, and combining them with preset events of interest, the problem of inaccurate noise levels in existing technologies has been solved, achieving a more stable and accurate noise level assessment.

CN116206629BActive Publication Date: 2026-08-25TP-LINK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310209017.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2026-08-25
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

Existing technologies struggle to obtain stable and accurate noise levels, and subjective human evaluation methods are heavily influenced by personal perceptions, resulting in low accuracy of noise classification information.

Method used

The noise level is determined by acquiring the noise and semantic features of the audio and video streams to be inspected, combined with preset events of interest. The noise features are derived from energy spectrum features extracted using traditional acoustic features and deep learning methods. The semantic features include the distribution of human voices and target objects. The noise level is calculated using neural networks and mapping rules.

Benefits of technology

It improves the stability and accuracy of noise levels, objectively reflects the impact of noise on audio and video streams, and adapts to the noise processing needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206629B_ABST
    Figure CN116206629B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of noise processing, and provides a noise processing method and device, electronic equipment and a computer readable storage medium, which comprises the following steps: obtaining noise features of a to-be-detected audio / video stream; obtaining content semantic features of the to-be-detected audio / video stream, wherein the content semantic features are determined according to a preset attention event; and determining a noise level of the to-be-detected audio / video stream according to the noise features and the content semantic features. Through the above method, the accuracy of the obtained noise level can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of noise treatment technology, and particularly relates to noise treatment methods, apparatus, electronic devices and computer-readable storage media. Background Technology

[0002] Audio noise and its related processing is a very fundamental field that has been extensively studied in academia and industry. However, the processing of audio noise usually involves noise detection, noise classification, and other similar tasks.

[0003] For example, noise detection can be performed on the audio to be tested to determine whether there is noise in the audio.

[0004] For example, after identifying the presence of noise in an audio file, further classification of the noise can determine whether it is steady-state noise or non-steady-state noise.

[0005] However, neither noise detection nor noise classification in audio can provide information about the extent to which noise affects people.

[0006] In order to determine the degree of impact (or classification information) of noise on people, different people can make subjective comparisons of the original audio and the audio after system processing and degradation, and obtain the corresponding average subjective opinion score (MOS). Then, the average value is calculated from each MOS score, and the average value is used as the final MOS score of the original audio.

[0007] However, the MOS scoring method, which is based on subjective human evaluation, is difficult to obtain stable and accurate noise classification information because it is strongly related to personal subjective feelings. Summary of the Invention

[0008] This application provides noise processing methods, apparatus, electronic devices, and computer-readable storage media, which can solve the problem that existing methods are difficult to obtain stable and accurate noise levels.

[0009] In a first aspect, embodiments of this application provide a noise processing method, including:

[0010] Obtain the noise characteristics of the audio and video streams to be inspected;

[0011] Obtain the semantic features of the audio and video stream to be inspected, wherein the semantic features are determined based on preset attention events;

[0012] The noise level of the audio and video stream to be inspected is determined based on the noise characteristics and the content semantic characteristics.

[0013] Secondly, embodiments of this application provide a noise processing apparatus, comprising:

[0014] The noise feature acquisition module is used to acquire the noise features of the audio and video streams to be inspected.

[0015] The content semantic feature acquisition module is used to acquire the content semantic features of the audio and video stream to be inspected, wherein the content semantic features are determined based on preset attention events;

[0016] The noise level determination module is used to determine the noise level of the audio and video stream to be inspected based on the noise characteristics and the content semantic characteristics.

[0017] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.

[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0019] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the method described in the first aspect above.

[0020] The beneficial effects of the embodiments in this application compared with the prior art are:

[0021] In this embodiment, noise features and content semantic features of the audio / video stream to be tested are obtained, and the noise level of the audio / video stream to be tested is determined based on the noise features and the content semantic features. Since the noise features objectively reflect the noise of the audio / video stream to be tested (i.e., the noise features are an objective reflection of the noise in the audio / video stream to be tested), and since the content semantic features are determined based on preset events of interest (i.e., the content semantic features are a subjective reflection of the noise in the audio / video stream to be tested), the noise level determined based on the noise features and the content semantic features can stably and accurately reflect the impact of noise on the audio / video stream to be tested, thereby improving the stability and accuracy of the obtained noise level. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0023] Figure 1 This is a schematic flowchart of a noise processing method provided in an embodiment of this application;

[0024] Figure 2This is a schematic flowchart of another noise processing method provided in an embodiment of this application;

[0025] Figure 3 This is a schematic diagram of the structure of a noise processing device provided in one embodiment of this application;

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0028] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0029] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0032] Example 1:

[0033] When it is necessary to determine the extent of noise's impact on people, the MOS (Must-Obtain) scoring method can be used. However, this method involves different people subjectively comparing the original corpus with a degraded version of the original corpus after systematic processing, resulting in multiple MOS scores. The average of these multiple MOS scores is then calculated as the final MOS score for the original corpus. Because the MOS score lacks quantitative basis and is a purely subjective qualitative measurement, and is greatly influenced by individual preferences, the accuracy of the noise classification information obtained is relatively low.

[0034] To improve the accuracy of the obtained noise classification information, embodiments of this application provide a noise processing method.

[0035] In this noise processing method, the user predefines the information they are interested in, such as human voices and vehicles. Before performing noise classification on the audio and video stream to be tested, the electronic device first acquires the noise characteristics of the audio and video stream to be tested and the semantic features of the content determined based on the aforementioned information of interest, and then determines the noise level of the audio and video stream to be tested based on the acquired features.

[0036] The noise processing method provided in the embodiments of this application is described below with reference to the accompanying drawings.

[0037] Figure 1 A flowchart of a noise processing method provided in an embodiment of this application is shown, which is applied to an electronic device, and is described in detail below:

[0038] Step S11: Obtain the noise characteristics of the audio and video stream to be inspected.

[0039] The audio and video stream to be inspected is a data stream that includes audio and video information, and the audio and video stream to be inspected can be an audio and video stream obtained from security equipment.

[0040] In this embodiment, the noise feature is a feature that can objectively reflect the noise of the audio / video stream under test. Specifically, the electronic device extracts audio information from the audio / video stream under test, and then extracts noise features from the audio information.

[0041] In some embodiments, considering that more accurate noise features can be extracted from stationary audio information, and that speech signals have short-term stationary characteristics, the audio information can be divided into multiple segments using a sliding window, and then the corresponding noise features can be extracted from each segment separately. The duration of each segment can be set to 10 milliseconds (ms) or 20 milliseconds (ms), etc., and is not limited here. Dividing the audio information into multiple segments facilitates the extraction of accurate noise features from each segment separately.

[0042] In some embodiments, considering that the stable speech signals corresponding to different scenarios may be different, the duration of the segments can be set according to the scenario corresponding to the audio / video stream to be detected, that is, the duration of the segments obtained after segmenting the audio information can be set according to the type of the scenario. Here, the scenario can be determined based on time and / or location, etc.

[0043] In this embodiment, noise features can be represented by traditional acoustic features such as zero-crossing rate, energy density, loudness, sharpness, dot intensity, and smoothness, or by energy spectrum features extracted by deep learning methods.

[0044] Traditional features typically consist of one or a few feature points. Their advantages are simplicity and low computational resource requirements, but their disadvantage is incomplete feature set. In contrast, energy spectrum features are extracted through deep learning rather than direct computation, making them more accurate and closer to human auditory perception in representing noise. However, their disadvantages include the need for larger computational resources and a certain amount of training data.

[0045] In this embodiment of the application, if the noise characteristics are represented by energy spectrum characteristics, and the audio / video stream to be tested is an audio / video stream collected by a security device, the energy spectrum characteristics of the audio / video stream to be tested can be obtained in the following ways:

[0046] A training dataset is constructed by acquiring audio and video streams collected by security equipment. This dataset is then used to train a neural network (such as RNNoise, RNN, Transformer, RCNN, etc.) to extract noise features. The trained neural network is then used to extract the energy spectrum features of the audio and video stream to be inspected, which are then used as the noise features of that stream.

[0047] In some embodiments, to improve the accuracy of the obtained noise features, the energy spectrum features extracted by the trained neural network are corrected, and the corrected energy spectrum features are used as the final noise features. Specifically, a person manually judges whether the extracted noise features are accurate by combining the video information in the audio and video stream. If they are inaccurate, some energy spectrum features are deleted or added to obtain corrected energy spectrum features, which are then used by the electronic device as the final noise features.

[0048] In some embodiments, considering that the frequency of sound waves audible to the human ear is limited, and only sound waves that can affect human hearing are judged as noise, in order to improve the speed of obtaining noise features and save the processing resources of electronic devices, only the energy distribution at a specified frequency point (i.e., the gain value at the specified frequency point) is extracted as the energy spectrum feature of this application embodiment. Here, the specified frequency point is a frequency point specified by the user, such as the frequency point where the sound wave frequency is audible to the human ear. Assuming that the audio information is divided into i segments, the energy spectrum feature can be F... i =[f 1,i ,...,f p,i ] indicates that, among which, f p,i Let p be the energy of the frequency point p of the i-th segment in the audio information.

[0049] Step S12: Obtain the semantic features of the audio and video stream to be inspected. The semantic features are determined based on preset attention events.

[0050] Specifically, video information is extracted from the audio and video stream to be inspected, and corresponding semantic features are extracted from the video information based on preset attention events. The attention information includes human voices and / or target objects, which can be human figures, vehicles, and / or beams of light, etc.

[0051] In some embodiments, if the audio information is divided into multiple segments and noise features are extracted from each segment separately, then before extracting the semantic features, the video information is first segmented according to the audio information segmentation rules, and then the corresponding semantic features are extracted from the segmented video information. It should be noted that since there is a one-to-one correspondence between the segmented audio information and the segmented video information, there is also a one-to-one correspondence between the noise features extracted from the segmented audio information and the semantic features extracted from the segmented video information. For example, assuming the audio information is divided into audio information 1 and audio information 2, and the noise features extracted from audio information 1 and audio information 2 are noise feature 1 and noise feature 2 respectively, then the video information is also divided into video information 1 and video information 2 according to the audio information segmentation rules (which include the order of segmentation and the duration of the segments, etc.). Assume that the semantic features of video information 1 and video information 2 are content semantic feature 1 and content semantic feature 2, respectively. Since audio information 1 corresponds to video information 1 and audio information 2 corresponds to video information 2, noise feature 1 corresponds to content semantic feature 1 and noise feature 2 corresponds to content semantic feature 2.

[0052] Step S13: Determine the noise level of the audio and video stream to be inspected based on the above noise characteristics and the above content semantic characteristics.

[0053] Specifically, noise features and content semantic features can be combined as a single entity, and a pre-defined correspondence between this combination and noise levels can be established. Once the noise features and content semantic features of the audio / video stream to be inspected are determined, the noise level corresponding to the audio / video stream is found based on the determined combination of noise features and content semantic features and the pre-defined correspondence. In other embodiments, weights corresponding to noise features and content semantic features can be pre-defined separately. Once the noise features and content semantic features of the audio / video stream to be inspected are determined, the noise level corresponding to the audio / video stream to be inspected is calculated based on the determined noise features and content semantic features and the weights corresponding to each feature.

[0054] In this embodiment, the noise level is used to indicate the degree of subjective and objective impact of noise in the audio / video stream under test on the stream. The noise level can be represented numerically; for example, a larger value indicates a lower noise level, meaning a lower degree of impact from the noise on the audio / video stream under test. The noise level can also be represented verbally, for example, categorized as severe, moderate, or slight, or as severe and perceptible to humans, severe but not easily perceptible to humans, perceptible but not severe, and not severe and not easily perceptible to humans.

[0055] In this embodiment, noise features and content semantic features of the audio / video stream to be tested are obtained, and the noise level of the audio / video stream to be tested is determined based on the noise features and the content semantic features. Since the noise features can objectively reflect the noise of the audio / video stream to be tested, that is, the noise features are an objective reflection of the noise of the audio / video stream to be tested, and since the content semantic features are determined based on preset attention events, that is, the content semantic features are a subjective reflection of the noise of the audio / video stream to be tested, the noise level determined based on the noise features and the content semantic features can more stably and accurately reflect the impact of noise on the audio / video stream to be tested, thereby improving the accuracy of the obtained noise level.

[0056] Example 2:

[0057] In some embodiments, the information of interest in this application includes human voices and / or target objects, which may be human figures, vehicles, and / or light beams, etc. Step S12 above includes:

[0058] A1. Obtain the human voice characteristics of the above-mentioned audio and video stream to be inspected, and calculate the overlap between the human voice and noise in the above-mentioned audio and video stream to be inspected based on the above-mentioned human voice characteristics and the above-mentioned noise characteristics.

[0059] And / or,

[0060] A2. Obtain the distribution characteristics of the target body in the above-mentioned audio and video streams to be inspected.

[0061] Specifically, before acquiring human voice features, the overlap between human voice and noise, and the distribution features of the target body, audio and video information can be extracted from the audio and video stream to be inspected. Then, the audio information can be divided into multiple segments, and the video information can be divided into corresponding segments based on the segment information of the audio information. Finally, human voice features, the overlap between human voice and noise can be extracted from each segment of audio information, and the distribution features of the target body can be extracted from each segment of video information.

[0062] In this embodiment, when the information of interest includes human voice, since both human voice and noise are sound information and sounds can influence each other, in addition to acquiring the human voice features corresponding to the human voice in the audio / video stream to be inspected, the overlap between the human voice and the noise is also calculated so that the noise level of the audio / video stream to be inspected can be determined more accurately. The overlap between human voice and noise reflects the number of time points where both human voice and noise coexist; for example, the more time points where both human voice and noise coexist, the higher the overlap between them.

[0063] In this embodiment, the human voice features can be represented by traditional acoustic features such as zero-crossing rate, energy density, loudness, sharpness, dot intensity, and smoothness, or by energy spectrum features extracted using deep learning methods. The method for extracting human voice features is the same as that for noise features described above, and will not be repeated here.

[0064] It should be noted that if the noise feature of this application embodiment reflects the energy distribution of noise at a specified frequency, then the human voice feature also reflects the energy distribution of human voice at a specified frequency.

[0065] In this embodiment, when the information of interest includes a target object, a pre-trained neural network (such as YOLO, DenseNet, UNet, ViT, etc.) capable of recognizing target objects is first used to identify the audio / video stream to be inspected. If the target object is identified, the time when the target object appears and the time when the target object does not appear are recorded to obtain the distribution characteristics of the target object. That is, in this embodiment, the distribution characteristics of the target object are used to indicate whether the target object appears in the audio / video stream to be inspected, and the time of its appearance. For example, if the number of target objects is equal to 1, two values ​​can be used to represent whether the target object appears in the frame of the video stream to be inspected. Specifically, if the target object appears in the frame at time 1, it is marked with "1" at time 1; if the target object does not appear in the frame at time 2, it is marked with "0" at time 2. Of course, if the number of target objects is greater than 1, more than 2 values ​​are needed to represent whether the target object appears in the frame of the video stream to be inspected, which will not be elaborated here.

[0066] In this embodiment, the voice features and the overlap between voice and noise in the audio / video stream to be inspected are obtained based on the information included in the attention information; alternatively, the distribution features of the target body in the audio / video stream to be inspected are obtained; or, the voice features, the overlap between voice and noise, and the distribution features of the target body in the audio / video stream to be inspected are obtained. Since the attention information influences the user's choice of noise processing strategy, and the features obtained above are all content semantic features corresponding to the user's pre-defined attention information, obtaining these features is beneficial for more accurately determining the noise level of the audio / video stream to be inspected.

[0067] Example 3:

[0068] In some embodiments, the noise feature is a noise energy spectrum feature, the human voice feature is a human voice energy spectrum feature, and step S13 includes:

[0069] B1. According to the preset first mapping rule, the energy of each frequency point in the human voice energy spectrum feature of the above-mentioned audio and video stream to be tested is mapped to a preset first value or a second value. The above-mentioned first mapping rule is used to determine the value mapped to the energy of each frequency point in the above-mentioned human voice energy spectrum feature.

[0070] Specifically, the aforementioned first mapping rule sets the energy (assumed to be the first energy) or energy range mapped to a first value, and sets the energy (assumed to be the second energy) or energy range mapped to a second value. During the mapping process, the energy of the current frequency point is compared with the first energy and the second energy (or the corresponding energy range) set by the first mapping rule. If the energy of the current frequency point equals the first energy, then the energy of the current frequency point is mapped to the first value; or, if the energy of the current frequency point is within the energy range corresponding to the first value, then the energy of the current frequency point is mapped to the first value. The aforementioned first and second values ​​can be set as needed, for example, the first value can be set to "1" and the second value to "0", etc., without limitation here.

[0071] B2. According to the preset second mapping rule, the energy of each frequency point in the noise energy spectrum feature of the above-mentioned audio and video stream to be tested is mapped to the first value or the second value. The second mapping rule is used to determine the value mapped to the energy of each frequency point in the above-mentioned noise energy spectrum feature.

[0072] Specifically, the second mapping rule sets the energy (assumed to be the third energy) or energy range mapped to the first value, and sets the energy (assumed to be the fourth energy) or energy range mapped to the second value. During mapping, the energy of the current frequency point is directly compared with the third energy (or energy range) or fourth energy (or energy range) set by the second mapping rule to determine whether the energy of the current frequency point is mapped to the first value or the second value.

[0073] It should be noted that, due to the difference in sensitivity of the human ear to noise and human voice, the first energy mentioned above is usually not equal to the third energy, and the second energy mentioned above is usually not equal to the fourth energy.

[0074] B3. Count the number of times that the above-mentioned human voice energy spectrum characteristics and the above-mentioned noise energy spectrum characteristics are mapped to the first value or the second value at the same frequency point, and calculate the overlap between human voice and noise in the above-mentioned audio and video streams to be tested based on the statistical results.

[0075] Specifically, assuming the human voice energy spectrum feature is the energy distribution of the human voice at n frequency points, and the noise energy spectrum feature is the energy distribution of the noise at those n frequency points, then for each frequency point, it is determined whether the value of the human voice energy spectrum feature mapped at that frequency point is the first value (or the second value), and similarly, whether the value of the noise energy spectrum feature mapped at that frequency point is the first value (or the second value). If they are equal, the count is incremented by 1. After determining the values ​​of both the human voice and noise energy spectrum features mapped at each frequency point, the accumulated count is taken as the final count. Based on this final count and the number of frequency points, the overlap between the human voice and the noise is calculated.

[0076] In this embodiment of the application, since for the same frequency point, if the value after mapping the human voice energy spectrum feature is equal to the value after mapping the noise energy spectrum feature, it indicates that the human voice and noise overlap at that frequency point. Therefore, by mapping the energy of the human voice energy spectrum feature and the noise energy spectrum feature at the frequency point to a first value or a second value, and then counting the number of people with the first value (or the second value) at the same frequency point, the overlap between human voice and noise can be calculated quickly and accurately.

[0077] In some embodiments, considering that the human ear has different sensitivities to sounds of different frequencies, in order to avoid the need for further processing of noise with too weak an impact, the energy in the human voice energy spectrum feature and the noise energy spectrum feature is subtracted by a preset energy, and then the corresponding mapping operation is performed. At this time, in step B1 above, mapping the energy of each frequency point in the human voice energy spectrum feature of the audio-visual stream to be tested to a preset first value or a second value includes:

[0078] B11. Subtract the energy of each frequency point in the human voice energy spectrum feature of the above-mentioned audio and video stream to be tested from the energy of the corresponding frequency point in the preset human voice energy spectrum threshold to obtain the subtracted human voice energy spectrum feature. The preset human voice energy spectrum threshold includes all frequency points in the above-mentioned human voice energy spectrum feature, and the energy corresponding to each frequency point of the preset human voice energy spectrum threshold is determined according to the sensitivity of the human ear to the sound of the above-mentioned frequency point.

[0079] It should be noted that since the human ear is usually sensitive to sound at different frequencies, the energy of the above-mentioned human voice energy spectrum threshold is usually different at different frequencies, thereby improving the accuracy of the human voice energy spectrum features obtained after subtraction.

[0080] B12. Map the energy of each frequency point in the above-subtracted human voice energy spectrum characteristics to a preset first or second value.

[0081] For example, if the energy spectrum characteristics of human voice are characterized using H i =[h 1,i ,...,h p,i ] indicates that h p,i Let H0 be the energy of the human voice in the i-th segment of the audio information at frequency p, and let H0 be the preset human voice energy spectrum threshold. 1,0 ,...,h p,0 The energy of each frequency point in the human voice energy spectrum feature is subtracted from the energy of the corresponding frequency point in the preset human voice energy spectrum threshold to obtain the subtracted human voice energy spectrum feature, i.e., "[h 1,i -h 1,0 ,...,h p,i -h p,0 If the first mapping rule above sets the energy range for mapping to the first value to be greater than 0, that is, for [h] 1,i -h 1,0 ,...,h p,i -h p,0 For any frequency point in the range, if the energy after subtracting the frequency point is greater than 0, then the energy of that frequency point is mapped to the first value; otherwise, if the energy after subtracting the frequency point is not greater than 0, then the energy of that frequency point is mapped to the second value.

[0082] Correspondingly, in step B2 above, mapping the energy of each frequency point in the noise energy spectrum characteristics of the audio / video stream to be tested to the first value or the second value includes:

[0083] B21. Subtract the energy of each frequency point in the noise energy spectrum feature of the above-mentioned audio and video stream to be tested from the energy of the corresponding frequency point in the preset noise energy spectrum threshold to obtain the noise energy spectrum feature after subtraction. The preset noise energy spectrum threshold includes all frequency points in the above-mentioned noise energy spectrum feature, and the energy corresponding to each frequency point of the preset noise energy spectrum threshold is determined according to the degree of influence of noise on human voice.

[0084] It should be noted that since the degree of influence of noise on human voice is usually different at different frequencies, the energy of the noise energy spectrum thresholds mentioned above is usually different at different frequencies, thereby improving the accuracy of the noise energy spectrum characteristics obtained after subtraction.

[0085] B22. Map the energy of each frequency point in the noise energy spectrum characteristics after the above difference to the first value or the second value mentioned above.

[0086] For example, if the noise energy spectrum characteristics are adopted using F i =[f 1,i ,...,f p,i ] indicates that, among which, f p,i Let F0 be the energy of the noise in the i-th segment of the audio information at frequency p, and let F0 be the preset noise energy spectrum threshold. 1,0 ,...,f p,0 The energy of each frequency point in the noise energy spectrum feature is subtracted from the energy of the corresponding frequency point in the preset noise energy spectrum threshold to obtain the noise energy spectrum feature after subtraction, i.e., "[f 1,i -f 1,0 ,...,f p,i -f p,0 If the second mapping rule above sets the energy range for mapping to the first value to be greater than 0, that is, for [f] 1,i -f 1,0 ,...,f p,i -f p,0 For any frequency point in the range, if the energy after subtracting the frequency point is greater than 0, then the energy of that frequency point is mapped to the first value; otherwise, if the energy after subtracting the frequency point is not greater than 0, then the energy of that frequency point is mapped to the second value.

[0087] In this embodiment, since the noise energy spectrum thresholds corresponding to human voice and noise take into account the sensitivity of the human ear to sounds of different frequencies, after subtracting the preset energy from the energy in the human voice energy spectrum features and the noise energy spectrum features respectively, the corresponding mapping operation can be performed to avoid the need to process noise with too weak an impact, thereby effectively saving system resources.

[0088] Example 4:

[0089] In some embodiments, when determining the noise level of the audio / video stream to be inspected based on the weights corresponding to pre-defined noise features and content semantic features, step S13 includes:

[0090] C1. Match the noise base to be matched with the standard noise base in the preset noise base library, wherein the noise base to be matched includes the noise features and the content semantic features, and any two of the standard noise bases include different noise features and / or different content semantic features.

[0091] Specifically, a noise base library is pre-constructed, which includes multiple standard noise bases. Each standard noise base includes two parameters: noise features and content semantic features. Any two standard noise bases have different values ​​under the noise features and / or under the content semantic features, that is, each standard noise base is different.

[0092] When matching a noise base to be matched with a standard noise base in a preset noise base library, the noise features of the noise base to be matched can be compared with the noise features of the standard noise base first. If they are the same, the content semantic features of the noise base to be matched can be compared with the content semantic features of the standard noise base. If they are the same, the noise base to be matched is determined to be a match with the standard noise base; otherwise, the noise base to be matched is determined not to be a match with the standard noise base. Since noise features are the features that users are more concerned about, matching noise features first and then matching content semantic features can find a matching standard noise base more quickly.

[0093] In some embodiments, assuming that the semantic features of the content include human voice features, the overlap between human voice and noise, and the distribution features of the target object, the preset noise library can be constructed in the following ways:

[0094] (1) Establish the entry format for the noise base library: number i, noise energy spectrum characteristics F i The degree of overlap between noise and human voice (V) i Is there a target object's screen event E? i Then the formula for the noise base library is expressed as:

[0095] NS = {n1, n2, ..., n i ,...}

[0096] Where NS is the noise base library, n i Let be the i-th standard noise basis in the noise basis library, and:

[0097] n i =[i,F i V i E i ]

[0098] Among them, F i V is a 1*p matrix, where p is an adjustable preset value. i ∈[0,1], when there is only one target, the above E i It can take the value 1 or 0, for example, when E i A value of 1 indicates the presence of a target object's screen event, while when E... i A value of 0 indicates that there are no screen events for the target object.

[0099] (2) Create datasets for the application scenarios of electronic devices. For example, if the application scenario of electronic devices is security, the audio and video streams collected by surveillance cameras can be obtained, and a dataset can be constructed based on the obtained audio and video streams. Appropriate noise features NS can be extracted from the self-built dataset. m ={m1,m2,...,m j ,...}, where NS m Let m be the set of noise features of the m-th segment of audio information (the length of the segment can be adjusted as needed). j Let j be the j-th noise feature, and:

[0100] m j =[j,F j V j E j ]

[0101] (3) Traverse NS m All noise features are compared with all standard noise bases in the noise base library for feature similarity. If there is a noise base with feature similarity within the threshold range (i.e., the same type), the current noise feature is assigned to that standard noise base, and the human voice overlap and video events of the noise feature are recorded in the corresponding entries of that standard noise base. Otherwise, the noise feature is added to the noise base library as a new standard noise base.

[0102] (4) The newly added noise features (such as noise energy spectrum features) in the noise base library are converted into audio files and manually reviewed in conjunction with video. This involves manually verifying whether the extracted noise features are indeed the features corresponding to the noise. If not, the noise base corresponding to the invalid noise features is removed. After manual review, the final structure of the standard noise base is as follows:

[0103]

[0104] That is, for the i-th standard noise basis n i Its noise characteristics are F i There are k values ​​for the overlap between human voice and noise, and k values ​​for video events, where k is a natural number greater than 0.

[0105] Since a single audio or video clip may contain multiple human voices and / or multiple types of visual events, combining noise features with multiple possible semantic features can improve the accuracy of the resulting standard noise base.

[0106] C2. If a target noise base is matched from the above noise base library, the weight matrix corresponding to the target noise base is determined according to the target noise base and the preset mapping relationship. The target noise base is a standard noise base that matches the noise base to be matched. The mapping relationship includes the correspondence between each of the standard noise bases and the weight matrix. The weight matrix is ​​used to record the weights corresponding to the noise features and / or content semantic features included in the standard noise base.

[0107] Specifically, the correspondence between each standard noise basis and the weight matrix in the noise basis library is pre-defined to obtain the above mapping relationship.

[0108] In this embodiment, the weights in the weight matrix are related to the information included in the standard noise basis. For example, if the standard noise basis includes noise features and content semantic features, then the weight matrix includes the weights corresponding to the noise features and the weights corresponding to the content semantic features. Furthermore, if the content semantic features include multiple values, such as W screen events, then the weights corresponding to the content semantic features also include W.

[0109] It should be noted that if no target noise base is found in the noise base library, the noise base to be matched is added to the noise base library as a standard noise base in the library.

[0110] C3. Determine the noise level of the audio and video streams to be tested based on the noise base to be matched and the weight matrix.

[0111] Specifically, each feature in the noise base to be matched is multiplied by its corresponding weight in the weight matrix. Then, based on the multiplication result and a preset level mapping table, the noise level of the audio and video stream to be detected is determined. The preset level mapping table is used to record the correspondence between the result of multiplying each feature (such as noise features and content semantic features) in each standard noise base with its corresponding weight and the noise level.

[0112] In this embodiment, since a noise base library is pre-constructed and the weight matrix corresponding to each standard noise base in the noise base library is pre-determined, the corresponding noise base to be matched can be quickly determined after the noise features and semantic features of the audio and video stream to be tested are determined. Furthermore, it is possible to quickly determine from the noise base library whether there is a standard noise base that matches the noise base to be matched, and then quickly determine the noise level based on the standard noise base and the corresponding weight matrix.

[0113] In some embodiments, the correspondence between each standard noise base in the noise base library and the weight matrix is ​​predetermined. Before determining the weight matrix corresponding to the target noise base based on the target noise base and the predetermined mapping relationship in step C2, the method further includes:

[0114] D1. Obtain the subjective opinion score of the audio and video stream corresponding to the above standard noise basis, wherein the scene corresponding to the audio and video stream corresponding to the above standard noise basis is the same as the scene corresponding to the audio and video stream to be tested.

[0115] In this embodiment, considering that users' subjective perception of noise in audio and video streams in different scenarios may vary—for example, users have lower noise tolerance in quiet work environments and higher noise tolerance in bustling street markets—it is advisable to pre-obtain audio and video streams with a standard noise base corresponding to the same scenario as the audio and video stream to be tested, and then obtain the subjective opinion score of that audio and video stream. This ensures the accuracy of the weight matrix subsequently determined based on the obtained subjective opinion score. In other words, in this embodiment, different noise base libraries are selected according to different scenarios corresponding to the audio and video stream to be tested to ensure the accuracy of the weight matrix subsequently determined based on the standard noise base in the selected noise base library.

[0116] The subjective opinion score can be obtained through the following methods:

[0117] When the scene of the audio and video stream to be inspected is a security scene, multiple audio and video streams in the security scene are collected, and then multiple users score the noise of the multiple audio and video streams respectively to obtain the subjective opinion score of each user for the noise of each audio and video stream. The average score of the subjective opinion score corresponding to each audio and video stream is taken as the subjective opinion score corresponding to that audio and video stream.

[0118] Because the subjective opinion score for the audio and video stream is obtained by directly scoring the noise of the audio and video stream by multiple users, instead of first extracting the audio information from the audio and video stream and then having multiple users score the noise of the extracted audio information, the influence of the video image on the user's perception of audio noise is preserved, thereby improving the stability and accuracy of the obtained subjective opinion score.

[0119] D2. Determine the weight matrix corresponding to the standard noise basis based on the above standard noise basis and the above subjective opinions.

[0120] Specifically, a deep learning algorithm can be used to learn the standard noise base and subjective opinion score to obtain the weights corresponding to each parameter in the standard noise base, and then obtain a weight matrix composed of the weights corresponding to each parameter (e.g., the weights corresponding to noise features and the weights corresponding to content semantic features).

[0121] D3. Based on the above standard noise basis and the above weight matrix, determine the above-preset mapping relationship.

[0122] Specifically, since each standard noise basis is different, the weight matrices corresponding to each standard noise basis are also different. For example, if standard noise basis A and standard noise basis B have the same noise features, but at least one parameter in the content semantic features is different, such as the human voice features of standard noise basis A and standard noise basis B being different, then standard noise basis A and standard noise basis B are different standard noise bases, and standard noise basis A and standard noise basis B correspond to different weight matrices.

[0123] In this embodiment of the application, since the influence of the scene on the subjective opinion score is considered when determining the weight matrix corresponding to the standard noise basis, the accuracy of the determined mapping relationship between the standard noise basis and the weight matrix is ​​improved.

[0124] In some embodiments, after step S13 described above, the method further includes:

[0125] Based on the noise level mentioned above, select the corresponding processing strategy to process the audio and video streams to be tested.

[0126] In this embodiment, since the noise level is used to indicate the degree of subjective and objective impact of noise in the audio / video stream under test, after obtaining the audio / video stream under test, a corresponding processing strategy can be selected to process the stream, thereby improving the accuracy of the processing results. For example, suppose the noise level is divided into severe and perceptible to humans, severe but not easily perceptible to humans, perceptible but not severe, and not severe and not easily perceptible to humans. When the noise level is severe and perceptible to humans, strong noise reduction must be performed. In this case, the selected processing strategy is a strong noise reduction strategy. When the noise level is severe but not easily perceptible to humans, noise reduction can be omitted to save computing resources. In this case, the selected processing strategy is a no-noise-reduction strategy. When the noise level is perceptible but not severe, conservative noise reduction can be performed to avoid excessive noise reduction causing side effects, such as voice distortion. In this case, the selected processing strategy is a weak noise reduction strategy.

[0127] To more clearly describe the noise processing method provided in the embodiments of this application, the following description is in conjunction with the appendix. Figure 2 Further description.

[0128] refer to Figure 2 The noise processing method of this application is mainly divided into three parts: noise base library establishment, feature extraction and hierarchical output.

[0129] The noise base library construction process involves: first, establishing the entry format for the noise base library, that is, defining the information contained in each standard noise base; then, extracting the required features from the self-built dataset, such as noise features and semantic features. After obtaining each feature, the audio information converted from the noise features is verified by combining human hearing with video information to eliminate invalid noise bases and use the valid noise bases as the standard noise bases in the noise base library.

[0130] Feature extraction process: Audio information is extracted from the audio and video stream to be inspected and divided into appropriate audio segments; corresponding video segments are extracted from the audio and video stream based on the audio segments. For each audio segment, noise features, human voice features, and the overlap between human voice and noise are extracted. For each video segment, the distribution features of the target object are extracted. The noise features, human voice features, the overlap between human voice and noise, and the distribution features of the target object are combined to form the noise basis for matching.

[0131] The hierarchical output process is as follows: The noise base to be matched is matched with the standard noise base in the noise base library. If a standard noise base is obtained (i.e., the target noise base is matched), the weight matrix corresponding to the target noise base is found, and the noise level of the audio and video stream to be tested is calculated based on the noise base to be matched and the weight matrix.

[0132] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0133] Example 5:

[0134] Corresponding to the noise processing method described in the above embodiments, Figure 3 A structural block diagram of a noise processing apparatus provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0135] Reference Figure 3 The noise processing device 3 is applied to electronic devices and includes: a noise feature acquisition module 31, a content semantic feature acquisition module 32, and a noise level determination module 33. Wherein:

[0136] The noise feature acquisition module 31 is used to acquire the noise features of the audio and video stream to be inspected.

[0137] The content semantic feature acquisition module 32 is used to acquire the content semantic features of the above-mentioned audio and video stream to be inspected. The content semantic features are determined according to preset attention events.

[0138] The noise level determination module 33 is used to determine the noise level of the audio and video stream to be inspected based on the noise characteristics and the semantic features of the content.

[0139] In this embodiment, noise features and content semantic features of the audio / video stream to be tested are obtained, and the noise level of the audio / video stream to be tested is determined based on the noise features and the content semantic features. Since the noise features can objectively reflect the noise of the audio / video stream to be tested, that is, the noise features are an objective reflection of the noise of the audio / video stream to be tested, and since the content semantic features are determined based on preset attention events, that is, the content semantic features are a subjective reflection of the noise of the audio / video stream to be tested, the noise level determined based on the noise features and the content semantic features can more accurately reflect the impact of noise on the audio / video stream to be tested, thereby improving the accuracy of the obtained noise level.

[0140] In some embodiments, the content semantic feature acquisition module 32 includes:

[0141] The human voice feature acquisition module is used to acquire the human voice features of the audio and video stream to be inspected, and to calculate the overlap between the human voice and noise in the audio and video stream to be inspected based on the human voice features and the noise features.

[0142] And / or,

[0143] The target body distribution feature acquisition module is used to acquire the distribution features of the target body in the above-mentioned audio and video stream to be inspected.

[0144] In some embodiments, the noise feature is a noise energy spectrum feature, the human voice feature is a human voice energy spectrum feature, and the human voice feature acquisition module, when calculating the overlap between the human voice and noise in the audio / video stream to be inspected based on the human voice feature and the noise feature, is specifically used for:

[0145] According to the preset first mapping rule, the energy of each frequency point in the human voice energy spectrum feature of the above-mentioned audio and video stream to be tested is mapped to a preset first value or a second value. The first mapping rule is used to determine the value mapped to the energy of each frequency point in the above-mentioned human voice energy spectrum feature.

[0146] According to the preset second mapping rule, the energy of each frequency point in the noise energy spectrum feature of the above-mentioned audio and video stream to be tested is mapped to the first value or the second value to obtain the second mapping result. The second mapping rule is used to determine the value mapped to the energy of each frequency point in the above-mentioned noise energy spectrum feature.

[0147] The number of times that the above-mentioned human voice energy spectrum characteristics and the above-mentioned noise energy spectrum characteristics are both mapped to the first value or the second value at the same frequency point is counted, and the overlap between human voice and noise in the above-mentioned audio and video streams to be tested is calculated based on the statistical results.

[0148] In some embodiments, mapping the energy of each frequency point in the human voice energy spectrum features of the audio / video stream to be detected to a preset first value or a second value includes:

[0149] The energy of each frequency point in the human voice energy spectrum feature of the above-mentioned audio and video stream to be tested is subtracted from the energy of the corresponding frequency point in the preset human voice energy spectrum threshold to obtain the subtracted human voice energy spectrum feature. The preset human voice energy spectrum threshold includes all frequency points in the above-mentioned human voice energy spectrum feature, and the energy corresponding to each frequency point of the preset human voice energy spectrum threshold is determined according to the sensitivity of the human ear to the sound of the above-mentioned frequency point.

[0150] The energy of each frequency point in the above-subtracted human voice energy spectrum characteristics is mapped to a preset first or second value.

[0151] Correspondingly, mapping the energy of each frequency point in the noise energy spectrum characteristics of the audio / video stream to be tested to the first or second value includes:

[0152] The energy of each frequency point in the noise energy spectrum feature of the above-mentioned audio and video stream to be tested is subtracted from the energy of the corresponding frequency point in the preset noise energy spectrum threshold to obtain the noise energy spectrum feature after subtraction. The preset noise energy spectrum threshold includes all frequency points in the above-mentioned noise energy spectrum feature, and the energy corresponding to each frequency point of the preset noise energy spectrum threshold is determined according to the degree of influence of noise on human voice.

[0153] The energy of each frequency point in the noise energy spectrum characteristics after the above subtraction is mapped to the first value or the second value mentioned above.

[0154] In some embodiments, the noise level determination module 33 includes:

[0155] The noise base matching unit is used to match the noise base to be matched with the standard noise base in the preset noise base library. The noise base to be matched includes the noise features and the content semantic features. Any two of the standard noise bases include different noise features and / or different content semantic features.

[0156] The weight matrix determination unit is used to determine the weight matrix corresponding to the target noise base according to the target noise base and a preset mapping relationship if a target noise base is matched from the above noise base library. The target noise base is a standard noise base that matches the noise base to be matched. The mapping relationship includes the correspondence between each of the standard noise bases and the weight matrix. The weight matrix is ​​used to record the weights corresponding to the noise features and / or content semantic features included in the standard noise base.

[0157] The noise level calculation unit is used to determine the noise level of the audio and video stream to be tested based on the noise base to be matched and the weight matrix.

[0158] In some embodiments, the noise processing device 3 further includes:

[0159] The subjective opinion score acquisition module is used to acquire the subjective opinion score of the audio and video stream corresponding to the above-mentioned standard noise basis, wherein the scene corresponding to the audio and video stream corresponding to the above-mentioned standard noise basis is the same as the scene corresponding to the audio and video stream to be tested.

[0160] The weight matrix determination module is used to determine the weight matrix corresponding to the standard noise basis based on the standard noise basis and the subjective opinion.

[0161] The mapping relationship determination module is used to determine the preset mapping relationship based on the standard noise basis and the weight matrix.

[0162] In some embodiments, the noise processing device 3 further includes:

[0163] The noise processing module is used to select the corresponding processing strategy based on the noise level to process the audio and video streams to be tested.

[0164] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0165] Example 6:

[0166] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 4 of this embodiment includes: at least one processor 40 ( Figure 4 The diagram shows only one processor, a memory 41, and a computer program 42 stored in the memory 41 and executable on the at least one processor 40, which, when executing the computer program 42, performs the steps in any of the above method embodiments.

[0167] The electronic device 4 can be a computing device such as a surveillance camera, desktop computer, laptop, handheld computer, or cloud server. This electronic device may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 4 and does not constitute a limitation on electronic device 4. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0168] The processor 40 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0169] In some embodiments, the memory 41 may be an internal storage unit of the electronic device 4, such as a hard disk or memory of the electronic device 4. In other embodiments, the memory 41 may be an external storage device of the electronic device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 4. Furthermore, the memory 41 may include both internal and external storage units of the electronic device 4. The memory 41 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 41 can also be used to temporarily store data that has been output or will be output.

[0170] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0171] This application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.

[0172] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0173] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.

[0174] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0175] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0176] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0178] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0179] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A noise processing method, characterized in that, include: Obtain the noise characteristics of the audio and video streams to be inspected; The content semantic features of the audio and video stream to be inspected are obtained. The content semantic features are determined based on preset attention events. The content semantic features are a subjective reflection of the noise in the audio and video stream to be inspected. The noise level of the audio and video stream to be inspected is determined based on the noise characteristics and the content semantic characteristics; The events of interest include human voices and / or target objects, and the acquisition of the semantic features of the content of the audio and video stream to be inspected includes: The voice features of the audio and video stream to be inspected are obtained, and the overlap between the voice and noise in the audio and video stream to be inspected is calculated based on the voice features and the noise features. And / or, Obtain the distribution characteristics of the target body in the audio and video stream to be inspected.

2. The noise processing method as described in claim 1, characterized in that, The noise feature is a noise energy spectrum feature, and the human voice feature is a human voice energy spectrum feature. The calculation of the overlap between the human voice and noise in the audio / video stream to be inspected based on the human voice feature and the noise feature includes: According to a preset first mapping rule, the energy of each frequency point in the human voice energy spectrum feature of the audio and video stream to be tested is mapped to a preset first value or a second value. The first mapping rule is used to determine the value mapped to the energy of each frequency point in the human voice energy spectrum feature. According to the preset second mapping rule, the energy of each frequency point in the noise energy spectrum feature of the audio and video stream to be tested is mapped to the first value or the second value. The second mapping rule is used to determine the value mapped to the energy of each frequency point in the noise energy spectrum feature. The number of times that the human voice energy spectrum feature and the noise energy spectrum feature are both mapped to a first value or a second value at the same frequency point is counted, and the overlap between the human voice and noise in the audio and video stream to be tested is calculated based on the statistical results.

3. The noise processing method as described in claim 2, characterized in that, The step of mapping the energy of each frequency point in the human voice energy spectrum feature of the audio / video stream to be detected to a preset first value or a second value includes: The energy of each frequency point in the human voice energy spectrum feature of the audio and video stream to be tested is subtracted from the energy of the corresponding frequency point in the preset human voice energy spectrum threshold to obtain the subtracted human voice energy spectrum feature. The preset human voice energy spectrum threshold includes all frequency points in the human voice energy spectrum feature, and the energy corresponding to each frequency point of the preset human voice energy spectrum threshold is determined according to the sensitivity of the human ear to the sound of the frequency point. Map the energy of each frequency point in the subtracted human voice energy spectrum feature to a preset first or second value; The step of mapping the energy of each frequency point in the noise energy spectrum characteristics of the audio / video stream to be tested to the first value or the second value includes: The energy of each frequency point in the noise energy spectrum feature of the audio and video stream to be tested is subtracted from the energy of the corresponding frequency point in the preset noise energy spectrum threshold to obtain the noise energy spectrum feature after subtraction. The preset noise energy spectrum threshold includes all frequency points in the noise energy spectrum feature, and the energy corresponding to each frequency point of the preset noise energy spectrum threshold is determined according to the degree of influence of noise on human voice. The energy of each frequency point in the subtracted noise energy spectrum characteristics is mapped to the first value or the second value.

4. The noise processing method as described in claim 1, characterized in that, Determining the noise level of the audio / video stream to be inspected based on the noise features and the content semantic features includes: The noise base to be matched is matched with the standard noise base in the preset noise base library, wherein the noise base to be matched includes the noise features and the content semantic features, and any two standard noise bases include different noise features and / or different content semantic features; If a target noise base is matched from the noise base library, the weight matrix corresponding to the target noise base is determined according to the target noise base and the preset mapping relationship. The target noise base is a standard noise base that matches the noise base to be matched. The mapping relationship includes the correspondence between each standard noise base and the weight matrix. The weight matrix is ​​used to record the weights corresponding to the noise features and / or content semantic features included in the standard noise base. The noise level of the audio and video stream to be tested is determined based on the noise base to be matched and the weight matrix.

5. The noise treatment method as described in claim 4, characterized in that, Before determining the weight matrix corresponding to the target noise basis based on the target noise basis and the preset mapping relationship, the method further includes: Obtain the subjective opinion score of the audio and video stream corresponding to the standard noise base, wherein the scene corresponding to the audio and video stream corresponding to the standard noise base is the same as the scene corresponding to the audio and video stream to be tested; Determine the weight matrix corresponding to the standard noise base based on the standard noise base and the subjective opinion score; The preset mapping relationship is determined based on the standard noise basis and the weight matrix.

6. The noise processing method according to any one of claims 1 to 5, characterized in that, After determining the noise level of the audio / video stream to be inspected based on the noise features and the content semantic features, the method further includes: The corresponding processing strategy is selected based on the noise level to process the audio and video stream to be tested.

7. A noise treatment device, characterized in that, include: The noise feature acquisition module is used to acquire the noise features of the audio and video streams to be inspected. The content semantic feature acquisition module is used to acquire the content semantic features of the audio and video stream to be inspected. The content semantic features are determined according to preset attention events and are a subjective reflection of the noise in the audio and video stream to be inspected. A noise level determination module is used to determine the noise level of the audio and video stream to be inspected based on the noise features and the content semantic features; The events of interest include human voices and / or target objects, and the content semantic feature acquisition module includes: The human voice feature acquisition module is used to acquire the human voice features of the audio and video stream to be inspected, and to calculate the overlap between the human voice and the noise in the audio and video stream to be inspected based on the human voice features and the noise features. And / or, The target body distribution feature acquisition module is used to acquire the distribution features of the target body in the audio and video stream to be inspected.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech recognition method and device, computer equipment and storage medium

    CN114141237A

  • Audio signal processing system and audio signal processing method

    US20120095755A1