Advertisement voice recognition method and device, computer equipment, storage medium and product

Through sliding window processing and Fourier transform, the target frequency point features are extracted, and the voiceprint feature vector is determined, which solves the problem of advertising voice recognition, realizes efficient recognition and classification, and improves robustness and processing speed.

CN119993199APending Publication Date: 2025-05-13浙江省市场监管发展研究中心(浙江省平台经济监测中心浙江省广告监测中心)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510002488.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently identify and classify advertising voice information, especially in cases where advertisements are mixed with other audio content in broadcast programs and lack of explicitly separated marks.

Method used

By sliding the target sound data, the target frequency point data is extracted using Fourier transform, the target frequency point features (frequency and intensity) are obtained, and the voiceprint feature vector is determined based on these features, and then advertising sound recognition is performed.

Benefits of technology

It improves the robustness and efficiency of advertising sound recognition, and can accurately extract sound data in different noise environments, reduce data scale, improve advertising matching effect, and improve processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993199A_ABST
    Figure CN119993199A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an advertisement voice recognition method and device, computer equipment, a storage medium and a product, and the advertisement voice recognition method comprises the steps: carrying out the sliding window processing of target voice data, and obtaining the window data of a plurality of target windows; the target sound data is acquired from broadcasting equipment; processing the window data based on Fourier transform to obtain target frequency point data of a plurality of target windows; obtaining a target frequency point feature of the target frequency point data, wherein the target frequency point feature comprises a target frequency point frequency and a target frequency point intensity; determining a voiceprint feature vector based on the maximum value of the target frequency point feature; and identifying the advertisement sound data based on the voiceprint feature vector and the target sound data. The problem of advertisement voice recognition can be solved, and the recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an advertising sound recognition method, device, computer equipment, storage medium and product. Background Art

[0002] Radio programs only contain sound information, and radio programs usually contain various types of audio content, such as dialogue, music, news reports, advertising broadcasts, etc. How to efficiently identify advertising information from sound information, be able to identify and classify advertising content, understand the promotional strategies of different brands or products in a specific time period, and then analyze consumer preferences and reactions; can also adjust their own marketing plans by monitoring the frequency, duration and content of competitors' advertisements; can also measure the actual broadcast of advertisements, including the number of broadcasts, time points, etc., to evaluate the effectiveness of advertising; can also be combined with the radio station's listening rate data to help estimate the size of the target audience reached by the advertisement.

[0003] Since radio programs are unstructured data, multiple audio contents are mixed with advertisements without clear separation marks. Moreover, advertisements may appear in various forms, including voice broadcasts, inserts in background music, celebrity endorsements, sponsor acknowledgments, etc. Each form has different characteristics, making it difficult to recognize advertisement sounds. Summary of the invention

[0004] In view of this, the present invention provides an advertisement sound recognition method, apparatus, computer equipment, storage medium and product to solve the problem of advertisement sound recognition and improve recognition efficiency.

[0005] In a first aspect, the present invention provides an advertising sound recognition method, which includes: performing sliding window processing on target sound data to obtain window data of multiple target windows; the target sound data is collected from a broadcasting device; processing the window data based on Fourier transform to obtain target frequency data of multiple target windows; obtaining target frequency features of the target frequency data, the target frequency features including target frequency frequency and target frequency intensity; determining a voiceprint feature vector based on a maximum value of the target frequency feature; and recognizing the advertising sound data based on the voiceprint feature vector and the target sound data.

[0006] In this implementation, the application only directly processes the audio band, and uses the frequency and intensity of the target frequency point to determine the voiceprint feature vector, where the frequency can reflect the pitch of the sound, and the intensity can reflect the energy of the sound, which helps to distinguish the loudness and importance of the sound. Advertising sound recognition based on this can improve robustness. Whether in a quiet room or a noisy public place, the voiceprint extraction method based on frequency and intensity can better extract sound data, greatly reducing the data size while improving the advertising matching effect and processing speed.

[0007] In an optional embodiment, sliding window processing is performed on the target sound data to obtain window data of multiple target windows, including: performing sliding window processing on the target sound data, and calculating the sound intensity data in each target window using a root mean square calculation method, wherein the sliding window width is a preset sampling point, and the sliding window coverage is a preset coverage rate.

[0008] In this implementation, by adopting sliding window processing, local patterns and trends in the data can be effectively captured, changes in window data can be analyzed more finely, non-stationary data can be adapted, complex problems can be simplified, model robustness can be enhanced, and it has high flexibility and ease of use.

[0009] In an optional implementation, before processing the window data based on Fourier transform to obtain target frequency point data of multiple target windows, the method also includes: grouping the target windows according to a first preset number; and for each group of target windows, binary encoding the sound intensity data using a coding method.

[0010] In this implementation, the binary encoding of the sound intensity data not only provides an efficient storage and transmission method, but also ensures the accuracy and security of the data.

[0011] In an optional implementation, processing the window data based on Fourier transform to obtain target frequency point data for multiple target windows includes: converting binary-coded sound intensity data from the time domain to the frequency domain based on Fourier transform to generate target frequency point data corresponding to each group of target windows.

[0012] In this implementation, data is converted from the time domain to the frequency domain based on Fourier transform, which can clearly separate different frequency contents, reduce the impact of background noise, facilitate subsequent analysis of target frequency point data, and improve recognition efficiency.

[0013] In an optional implementation, the target frequency feature is the target frequency frequency, and determining the voiceprint feature vector based on the maximum value of the target frequency feature includes: sorting the target frequency frequencies of the target frequency data; and obtaining a second preset number of target frequency data with the largest target frequency frequencies as the voiceprint feature vector.

[0014] In an optional embodiment, the target frequency feature is the target frequency intensity, and determining the voiceprint feature vector based on the maximum value of the target frequency feature includes: obtaining the maximum target frequency intensity in each target window, and obtaining the target position information of the maximum target frequency data; calculating the frequency difference based on the target frequency data and the corresponding target position information, the frequency difference including the adjacent frequency difference and the window position difference; sorting the frequency differences, and obtaining the third preset number of target frequency data with the largest difference as the voiceprint feature vector.

[0015] In a second aspect, the present invention provides an advertising sound recognition device, which includes: a first processing module, used to perform sliding window processing on target sound data to obtain window data of multiple target windows; the target sound data is collected from a broadcasting device; a second processing module, used to process the window data based on Fourier transform to obtain target frequency data of multiple target windows; an acquisition module, used to obtain target frequency features of the target frequency data, the target frequency features include target frequency frequency and target frequency intensity; a determination module, used to determine a voiceprint feature vector based on the maximum value of the target frequency feature; and an identification module, used to identify the advertising sound data based on the voiceprint feature vector and the target sound data.

[0016] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the advertising sound recognition method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the advertising sound recognition method of the first aspect or any corresponding embodiment thereof.

[0018] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the advertising sound recognition method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1is a flow chart of an advertisement sound recognition method according to an embodiment of the present invention;

[0021] Figure 2 is a flow chart of another advertising sound recognition method according to an embodiment of the present invention;

[0022] Figure 3 is a structural block diagram of an advertisement sound recognition device according to an embodiment of the present invention;

[0023] Figure 4 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0025] According to an embodiment of the present invention, an embodiment of an advertising sound recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0026] In this embodiment, a method for recognizing advertisement sound is provided. Figure 1 is a flow chart of an advertising sound recognition method according to an embodiment of the present invention. It should be noted that if there are substantially the same results, this embodiment is not limited to the above. Figure 1 The process sequence shown is limited. Figure 1 As shown, the process includes the following steps:

[0027] Step S101, performing sliding window processing on target sound data to obtain window data of multiple target windows.

[0028] The target sound data is collected from the broadcasting equipment. Specifically, the broadcasting sound information is collected from the broadcasting programs of the radio or tuner. The broadcasting programs include news programs and advertising programs. The method of the present application is used to identify the advertising programs in the broadcasting programs.

[0029] The target sound data is processed by sliding window method. Specifically, the sliding window width and sliding window coverage are set, and the target sound data is processed segment by segment according to the sliding window width to obtain multiple sound data segments and obtain window data of each target window.

[0030] For example, for the target sound data, if the sampling rate of the target sound data is 44.1kHz, that is, 44,100 samples are collected per second, and the sliding window processing is performed using a sliding window width of 2048 sampling points, the window data of each target window is approximately 46.4 milliseconds of sound data. If the sliding window processing is performed using a sliding window width of 4096 sampling points, the window data of each target window is approximately 92.9 milliseconds of sound data.

[0031] Step S102, processing the window data based on Fourier transform to obtain target frequency point data of multiple target windows.

[0032] The window data of each target window is converted from the time domain to the frequency domain using the Fourier transform method to obtain the target frequency point data of each target window.

[0033] Step S103, obtaining target frequency features of the target frequency data, where the target frequency features include target frequency frequency and target frequency intensity.

[0034] Among them, the frequency point intensity is the amplitude or power of a specific frequency component. It reflects the relative importance or contribution of the frequency component in the signal. The frequency point frequency refers to the frequency itself, that is, the number of times the signal completes periodic changes in a unit time.

[0035] Perform feature analysis on the target frequency data to determine the corresponding target frequency and target frequency intensity.

[0036] Step S104: determining a voiceprint feature vector based on the maximum value of the target frequency feature.

[0037] In one implementation, target frequency data corresponding to the largest preset number of target frequency frequencies are selected as voiceprint feature vectors. In another implementation, target frequency data corresponding to the largest preset number of target frequency intensities are selected as voiceprint feature vectors.

[0038] Step S105: identifying the advertisement sound data based on the voiceprint feature vector and the target sound data.

[0039] The entire target sound data is saved as a time vector to obtain the voiceprint comparison feature vector.

[0040] Among them, the time vector, also known as the time-frequency vector or time-frequency representation, is a mathematical tool used to describe the characteristics of a signal in two dimensions, time and frequency, combining information in the time domain and frequency domain.

[0041] The voiceprint comparison feature vector is compared with the voiceprint feature vector to identify the advertising sound data therein.

[0042] The advertising sound recognition method provided in this embodiment directly processes only the audio frequency band and determines the voiceprint feature vector using the target frequency point frequency and the target frequency point intensity, wherein the frequency can reflect the pitch of the sound, and the intensity can reflect the energy of the sound, which helps to distinguish the loudness and importance of the sound. Advertising sound recognition based on this can improve robustness. Whether in a quiet room or a noisy public place, the voiceprint extraction method based on frequency and intensity can better extract sound data, greatly reducing the data size while improving the advertising matching effect and processing speed.

[0043] In this embodiment, a method for recognizing advertisement sound is provided. Figure 2 is a flowchart of another advertising sound recognition method according to an embodiment of the present invention. It should be noted that if there are substantially the same results, this embodiment is not limited to the above method. Figure 2 The process sequence shown is limited. Figure 2 As shown, the process includes the following steps:

[0044] Step S201, performing sliding window processing on target sound data to obtain window data of multiple target windows.

[0045] The target sound data is collected from the broadcasting equipment.

[0046] Specifically, the above step S201 includes:

[0047] Step S2011, performing sliding window processing on the target sound data, and calculating the sound intensity data in each target window using a root mean square calculation method.

[0048] In one implementation, an overlapping sliding window method is used to perform sliding window processing on the target sound data.

[0049] Among them, overlapping sliding windows extract local information or features by moving a fixed-size window on the input data and allowing a certain overlap area between adjacent windows. It helps to capture patterns and trends in continuous data and increase the robustness of the model to position changes.

[0050] The sliding window width is set to the preset sampling points, and the sliding window coverage is set to the preset coverage, wherein the preset sampling points may be 2048 sampling points or 4096 sampling points, and the preset coverage may be 25% or 50%.

[0051] For example, the sliding window width is 2048 sampling points, and the sliding window coverage is 50%. When the target sound data is divided into sliding windows, 50% of the area between two adjacent target creations is overlapped.

[0052] For the target sound data in each target window, the root mean square of the instantaneous sound pressure value within the target window time is calculated by using the root mean square to obtain the sound intensity data representing the sound intensity level of the target sound data in the target window.

[0053] In one implementation, the sound intensity data is subjected to bandpass filtering to remove background noise.

[0054] In this implementation, by adopting sliding window processing, local patterns and trends in the data can be effectively captured, changes in window data can be analyzed more finely, non-stationary data can be adapted, complex problems can be simplified, model robustness can be enhanced, and it has high flexibility and ease of use.

[0055] Step S2012: group the target windows according to a preset number, and for each group of target windows, perform binary encoding on the sound intensity data using an encoding method.

[0056] The target window is divided into a plurality of window groups, and a plurality of sound intensity data in each group of data is binary-coded.

[0057] Exemplarily, the preset number is 64, and binary encoding is performed on every 64 sound intensity data.

[0058] In this implementation, the binary encoding of the sound intensity data not only provides an efficient storage and transmission method, but also ensures the accuracy and security of the data.

[0059] Step S202, processing the window data based on Fourier transform to obtain target frequency point data of multiple target windows.

[0060] The binary-coded sound intensity data is converted from the time domain to the frequency domain based on Fourier transform to generate the target frequency point data corresponding to each group of target windows.

[0061] In this implementation, data is converted from the time domain to the frequency domain based on Fourier transform, which can clearly separate different frequency contents, reduce the impact of background noise, facilitate subsequent analysis of target frequency point data, and improve recognition efficiency.

[0062] Step S203, obtaining target frequency point features of target frequency point data.

[0063] The target frequency feature includes the target frequency frequency and the target frequency intensity. For details, please refer to step S103, which will not be described in detail here.

[0064] Step S204: determining a voiceprint feature vector based on the maximum value of the target frequency feature.

[0065] In one implementation, the target frequency feature is a target frequency frequency.

[0066] Specifically, the above step S204 includes:

[0067] Step S2041, sorting the target frequency point frequencies of the target frequency point data.

[0068] Specifically, the frequencies of the target frequency points are sorted from large to small or the frequencies of the target frequency points are sorted from small to large.

[0069] Step S2042, obtaining a second preset number of target frequency point data with the largest target frequency point frequency as a voiceprint feature vector.

[0070] Specifically, exemplarily, the second preset number is 20, and the 20 target frequency point data with the largest target frequency points are used as binary voiceprint feature vectors.

[0071] In one implementation, the target frequency feature is the target frequency intensity.

[0072] Specifically, the above step S204 includes:

[0073] Step S2041, obtaining the maximum target frequency point intensity in each target window, and obtaining the target position information of the maximum target frequency point data.

[0074] Specifically, the resonance peak algorithm is used to obtain the maximum target frequency intensity in each target window.

[0075] Among them, the formant is the local intensity maximum in the speech signal.

[0076] Specifically, before calculating the maximum target frequency intensity, the target data of the target window is preprocessed to remove any linear or nonlinear long-term trends to ensure signal stability, and a high-pass filter is used to enhance high-frequency components, which can improve the accuracy of resonance peak detection.

[0077] The specific method may be to search for the local maximum point in the spectrum diagram, or to estimate the reflection coefficient by modeling the vocal tract as a linear system, and then to derive the resonance peak parameters from the reflection coefficient, etc., to determine the positions of all resonance peaks and their corresponding intensity values, to select the largest frame as the maximum target frequency intensity in the target window, and to determine the corresponding target position, wherein the target position is taken as the leftmost end.

[0078] The frequency peak envelope is constructed using the maximum target frequency intensity of each group of target windows.

[0079] The frequency peak envelope refers to a curve formed by connecting a series of local maximum values, i.e., peak values, in the frequency domain. It characterizes the intensity variation trend of the signal at different frequencies and can provide important information about the signal spectrum structure.

[0080] Step S2042: Calculate the frequency difference based on the target frequency data and the corresponding target position information.

[0081] The frequency point difference includes the adjacent frequency point difference and the window position difference.

[0082] The adjacent frequency point difference refers to calculating the numerical difference between each frequency point, ie, frequency component, and the frequency points on the left and right adjacent to it in a frequency domain.

[0083] Specifically, for the target frequency strength data S(f), where f represents the frequency index, the target frequency data f corresponding to the maximum target frequency strength is i , and the difference between the target frequency point data on the left and right is expressed as:

[0084] ΔS left (f i )=S(f i )-S(f i-1 ).

[0085] ΔS right (f i )=S(f i )-S(f i-1 ).

[0086] The window position difference is the difference between different window positions when using a sliding window for time domain or frequency domain analysis. It is used to measure the changes in the same time point or frequency point under different windows.

[0087] Specifically, we identify a sliding window function with time as a variable, and t is the time index. For two consecutive time windows W f (t 1 ) and W f (t 2 ),in,

[0088] ΔW f (t 1 ,t 2 )=|W f (t 2 )-W f (t 1 )|.

[0089] Identifies the target frequency intensity of window W at frequency f at time t.

[0090] Step S2043, sorting the frequency point differences, and obtaining a third preset number of target frequency point data with the largest differences as voiceprint feature vectors.

[0091] Sorting is performed from large to small according to the frequency difference or from small to large according to the target frequency. Specifically, the third preset number is 20, and 20 groups of target frequency data corresponding to the 20 frequency differences with the largest target frequency are used as binary voiceprint feature vectors.

[0092] Step S205: identifying the advertisement sound data based on the voiceprint feature vector and the target sound data.

[0093] The entire target sound data is saved as a time vector to obtain the voiceprint comparison feature vector.

[0094] Specifically, in one implementation, the target sound data is segmented into short time segments, and Fourier transform is applied to each segment to obtain a two-dimensional matrix, in which rows represent different positions on the time axis and columns correspond to various discrete frequencies.

[0095] In another implementation, the target sound data is decomposed using a series of wavelet basis functions, which may have good localization performance in both time and frequency.

[0096] The voiceprint comparison feature vector is compared with the voiceprint feature vector to identify the advertising sound data therein.

[0097] The advertising sound recognition method provided in this embodiment only directly processes the audio frequency band and determines the voiceprint feature vector using the target frequency point frequency and the target frequency point intensity, wherein the frequency can reflect the pitch of the sound, and the intensity can reflect the energy of the sound, which helps to distinguish the loudness and importance of the sound. Advertising sound recognition based on this can improve robustness. Whether in a quiet room or a noisy public place, the voiceprint extraction method based on frequency and intensity can better extract sound data, greatly reducing the data size while improving the advertising matching effect and processing speed.

[0098] In this embodiment, an advertising sound recognition device is also provided, which is used to implement the above embodiments and preferred implementations, and the descriptions that have been made will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0099] This embodiment provides an advertisement sound recognition device, such as Figure 3 As shown, including:

[0100] The first processing module 301 is used to perform sliding window processing on the target sound data to obtain window data of multiple target windows; the target sound data is collected from the broadcasting equipment.

[0101] The second processing module 302 is used to process the window data based on Fourier transform to obtain target frequency point data of multiple target windows.

[0102] The acquisition module 303 is used to acquire the target frequency point characteristics of the target frequency point data, where the target frequency point characteristics include the target frequency point frequency and the target frequency point intensity.

[0103] The determination module 304 is used to determine the voiceprint feature vector based on the maximum value of the target frequency point feature.

[0104] The recognition module 305 is used to recognize the advertisement sound data based on the voiceprint feature vector and the target sound data. The further functional description of each of the above modules is the same as that of the above corresponding embodiment, which will not be repeated here.

[0105] The advertising sound recognition device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0106] The embodiment of the present invention also provides a computer device having the above Figure 3 The advertising sound recognition device shown.

[0107] See also Figure 4 , Figure 4 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 4 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.

[0108] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0109] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0110] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0111] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0112] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.

[0113] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0114] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0115] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0116] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for recognizing advertisement sound, characterized in that: The method comprises: Performing sliding window processing on target sound data to obtain window data of multiple target windows; the target sound data is collected from broadcasting equipment; Processing the window data based on Fourier transform to obtain target frequency point data of multiple target windows; Acquire target frequency features of the target frequency data, wherein the target frequency features include target frequency frequency and target frequency intensity; Determine the voiceprint feature vector based on the maximum value of the target frequency feature; The advertisement sound data is identified based on the voiceprint feature vector and the target sound data.

2. The advertising sound recognition method according to claim 1, characterized in that: The step of performing sliding window processing on the target sound data to obtain window data of multiple target windows includes: The target sound data is subjected to sliding window processing, and the sound intensity data in each target window is calculated using a root mean square calculation method, wherein the sliding window width is a preset sampling point, and the sliding window coverage is a preset coverage rate.

3. The advertising sound recognition method according to claim 2, characterized in that: Before the processing of the window data based on Fourier transform to obtain target frequency point data of a plurality of target windows, the following further comprises: Grouping the target windows according to a first preset number; For each group of the target windows, the sound intensity data is binary-encoded using an encoding method.

4. The advertising sound recognition method according to claim 3, characterized in that: The processing of the window data based on Fourier transform to obtain target frequency point data of a plurality of target windows includes: The binary-coded sound intensity data is converted from the time domain to the frequency domain based on Fourier transform, so as to generate target frequency point data corresponding to each group of the target windows.

5. The advertising sound recognition method according to claim 1, characterized in that: The target frequency feature is the target frequency frequency, and determining the voiceprint feature vector based on the maximum value of the target frequency feature includes: Sorting the target frequency point frequencies of the target frequency point data; A second preset number of target frequency point data with the largest frequency of the target frequency point is obtained as the voiceprint feature vector.

6. The advertising sound recognition method according to claim 1, characterized in that: The target frequency feature is the target frequency intensity, and determining the voiceprint feature vector based on the maximum value of the target frequency feature includes: Obtaining the maximum target frequency point intensity in each of the target windows, and obtaining the target position information of the maximum target frequency point data; Calculate a frequency difference based on the target frequency data and the corresponding target position information, wherein the frequency difference includes an adjacent frequency difference and a window position difference; The frequency point differences are sorted, and a third preset number of target frequency point data with the largest differences are obtained as the voiceprint feature vectors.

7. An advertising sound recognition device, characterized in that: The device comprises: A first processing module is used to perform sliding window processing on target sound data to obtain window data of multiple target windows; the target sound data is collected from a broadcasting device; A second processing module is used to process the window data based on Fourier transform to obtain target frequency point data of multiple target windows; An acquisition module, used to acquire target frequency features of the target frequency data, wherein the target frequency features include target frequency frequency and target frequency intensity; A determination module, used to determine a voiceprint feature vector based on a maximum value of a target frequency feature; The recognition module is used to recognize the advertisement sound data based on the voiceprint feature vector and the target sound data.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the advertising sound recognition method according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the advertising sound recognition method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises computer instructions for causing a computer to execute the advertising sound recognition method according to any one of claims 1 to 6.