Video content recognition methods, devices, equipment, storage media, and software products

By combining video comments and descriptive text, and utilizing semantic analysis and audio-video fusion models to remove background noise, the problem of low accuracy in video information extraction due to background music is solved, achieving more accurate video content recognition.

CN122313344APending Publication Date: 2026-06-30CHENGDU TD TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU TD TECH LTD
Filing Date
2024-12-27
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

When background music is present in a video, the accuracy of audio recognition for extracting information is relatively low.

Method used

By acquiring video comment data and description text, a semantic analysis model is used to determine the volume of the background music. If the volume is too high, the video data and audio data are input into the audio-video fusion extraction model. The background noise is removed by combining the fusion information text and description text to obtain the noise-reduced frequency. The video content information is then determined based on the noise-reduced frequency and the video data.

Benefits of technology

It improves the accuracy of information extraction in the presence of background music, and achieves greater accuracy in extracting information from videos with loud background noise after noise reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122313344A_ABST
    Figure CN122313344A_ABST
Patent Text Reader

Abstract

This application provides a video content recognition method, apparatus, device, storage medium, and program product, belonging to the field of data processing technology. The method includes: acquiring a video to be recognized, comment data of the video to be recognized, and descriptive text, wherein the video to be recognized consists of video data and audio data; inputting the comment data into a semantic analysis model to obtain a judgment result of the background music volume; if the volume judgment result indicates a large volume, inputting the video data and audio data into an audio-video fusion extraction model to obtain fusion information text output by the audio-video fusion extraction model; using the fusion information text and the descriptive text to remove background noise from the audio data to obtain a noise-reduced frequency; and determining video content information based on the noise-reduced frequency and the video data. This method solves the problem of low accuracy in audio recognition and information extraction when background music is present.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a video content recognition method, apparatus, device, storage medium, and program product. Background Technology

[0002] Current information collection processes often require extracting information from videos. When it is necessary to organize the information in a video, the video content needs to be identified.

[0003] Currently, related technologies typically employ text recognition and audio recognition methods to extract information from videos.

[0004] However, the inventors discovered that the relevant technology has at least the following technical problems: videos usually have background music, and in this case, the accuracy of extracting information from the video using audio recognition will decrease. Summary of the Invention

[0005] This application provides video content recognition methods, apparatus, devices, storage media, and program products to solve the problem of low accuracy in audio recognition and information extraction when there is background music.

[0006] In a first aspect, embodiments of this application provide a video content recognition method, comprising: acquiring a video to be recognized, comment data of the video to be recognized, and descriptive text, wherein the video to be recognized consists of video data and audio data; inputting the comment data into a semantic analysis model to obtain a judgment result of the volume of background music; if the volume judgment result is that the volume is large, then inputting the video data and audio data into an audio-video fusion extraction model to obtain fusion information text output by the audio-video fusion extraction model; using the fusion information text and descriptive text to remove background noise from the audio data to obtain a noise-reduced frequency; and determining video content information based on the noise-reduced frequency and video data.

[0007] Secondly, embodiments of this application provide a video content recognition device, comprising: a data acquisition module for acquiring a video to be recognized, comment data of the video to be recognized, and descriptive text, wherein the video to be recognized consists of video data and audio data; a volume judgment module for inputting comment data into a semantic analysis model to obtain a judgment result on the volume of background music; a text acquisition module for inputting video data and audio data into an audio-video fusion extraction model if the volume judgment result indicates that the volume is large, to obtain fusion information text output by the audio-video fusion extraction model; a background noise removal module for removing background noise from the audio data using the fusion information text and descriptive text to obtain a noise-reduced frequency; and a content determination module for determining video content information based on the noise-reduced frequency and video data.

[0008] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0010] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0011] The video content recognition method, apparatus, device, storage medium, and program product provided in this application embodiment utilize the evaluation of loud background noise in video comments to determine whether the background music volume is too loud. When the background music volume is loud, fusion information is extracted from video data and audio data. The fusion information and descriptive text are used to remove background noise from the audio data to obtain a noise-reduced frequency. Based on the noise-reduced frequency and video data, video content information is determined. This achieves noise reduction of videos with loud background noise before extracting information from the video, increasing the accuracy of information extraction. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0013] Figure 1 A schematic diagram illustrating a scenario for the video content recognition method provided in this application;

[0014] Figure 2 A flowchart illustrating the video content recognition method provided in this application embodiment;

[0015] Figure 3 This is a schematic diagram of the structure of the video content recognition device provided in the embodiments of this application;

[0016] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0017] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0018] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0019] Currently, extracting information from videos has become an important task in the information gathering process, and organizing this information requires effective identification of video content. Mainstream technologies currently primarily extract information from videos by recognizing text within video images and employing audio recognition techniques.

[0020] However, excessive background noise in the video may affect the extraction of video content.

[0021] To address the aforementioned technical problems, the inventors propose the following technical concept: By combining video comments to determine whether the background noise is too loud, if the background noise is too loud, the audio-video fusion extraction model extracts fusion information, and by combining audio data, fusion information text, and descriptive text, the background noise in the audio data is removed to obtain the noise-reduced frequency. The video content information is then determined by combining the noise-reduced frequency and the video data.

[0022] This application is applied to scenarios involving video content recognition. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0023] Figure 1 This is a schematic diagram illustrating a scenario for the video content recognition method provided in this application. Figure 1 In this scenario, the components include: terminal device 101 and server 102.

[0024] In the specific implementation process, the terminal device 101 may include a computer, server, tablet, mobile phone, PDA (Personal Digital Assistant), and laptop, etc., which can input data.

[0025] Server 102 can be implemented using a single server or a cluster of multiple servers with more powerful processing capabilities and higher security. Where possible, it can also be replaced by a computer or laptop with strong computing power.

[0026] The connection between server 102 and terminal device 101 can be either wired or wireless.

[0027] Terminal device 101 is used to instruct server 102 to start acquiring the video to be identified, the comment data of the video to be identified, and the descriptive text, and to perform the video content recognition process.

[0028] It is understood that the scenarios illustrated in the embodiments of this application do not constitute a specific limitation on the video content recognition method. In other feasible embodiments of this application, the above scenarios may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components, which can be determined according to the actual application scenario and are not limited here. Figure 1 The scenario shown can be implemented by hardware, software, or a combination of both.

[0029] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0030] Figure 2 This is a flowchart illustrating the video content recognition method provided in an embodiment of this application. The execution entity of this embodiment may be... Figure 1 Server 102 in this embodiment is not particularly limited in this respect. Figure 2 As shown, the method includes:

[0031] S201: Obtain the video to be identified, the comment data of the video to be identified, and the description text, wherein the video to be identified consists of video data and audio data.

[0032] In this step, the video to be identified, its comment data, and descriptive text can be obtained from websites, software, and databases.

[0033] The descriptive text includes titles and summaries, while the comment data includes comments and bullet comments.

[0034] S202: Input the comment data into the semantic analysis model to obtain the result of judging the volume of the background music.

[0035] In this step, the semantic analysis model can be trained by staff using pre-labeled data, or it can be trained using the BERT model.

[0036] S203: If the volume judgment result is that the volume is too large, then the video data and audio data are input into the audio-video fusion extraction model to obtain the fusion information text output by the audio-video fusion extraction model.

[0037] In this step, the audio-video fusion extraction model can be composed of a video encoding network, an audio encoding network, a fusion encoding network, and a fusion decoding network. Video data is input into the video encoding network, audio data is input into the audio encoding network, the data output from the intermediate layer of the video encoding network and the audio encoding network is input into the fusion encoding network, and the output content of the fusion encoding network is input into the fusion decoding network to obtain the fused information text output by the fusion encoding network.

[0038] S204: By merging information text and description text, background noise is removed from the audio data to obtain noise-reduced audio.

[0039] This step includes removing descriptive text from the merged information text.

[0040] S205: Determine video content information based on noise-reduced audio and video data.

[0041] This step may include re-inputting the noise-reduced audio and video data into the audio-video fusion extraction model to obtain the video content information output by the model. Alternatively, it may include identifying the noise-reduced audio to obtain the corresponding first text, identifying the video data to obtain the corresponding second text, and combining the first and second texts to obtain the video content information.

[0042] Combining the first text and the second text can include deduplicating the first text and the second text and then superimposing them to obtain the video content information.

[0043] As can be seen from the description of the above embodiments, the embodiments of this disclosure utilize the evaluation of loud background noise in video comments to determine whether the background music volume is too loud. When the background music volume is loud, fusion information is extracted from video data and audio data. The fusion information and descriptive text are used to remove background noise from the audio data to obtain a noise-reduced frequency. Based on the noise-reduced frequency and video data, video content information is determined. This achieves noise reduction of videos with loud background noise before extracting information from the video, increasing the accuracy of information extraction.

[0044] In one possible implementation, step S204 above involves fusing information text and description text to remove background noise from audio data to obtain noise-reduced audio, including steps S2041 to S2045.

[0045] S2041: Use integrated information text and descriptive text to determine the text related to the video theme.

[0046] This step may include deduplicating the fused information text and description text to obtain video-themed text, or inputting the fused information and description text into a pre-trained large language model to obtain video-themed text output by the large language model.

[0047] S2042: Determine the audio text based on the audio data.

[0048] This step may include using an audio recognition algorithm or audio recognition model to recognize speech in the audio data to obtain audio text.

[0049] S2043: Remove video-related text from the audio text to obtain the background audio text.

[0050] This step may include calculating the similarity between each field in the audio text and the video topic-related text, and deleting fields with similarity greater than a preset similarity threshold from the calculated audio text to obtain background audio text; or it may include inputting the audio text and the video topic-related text into a pre-trained similar text removal large language model to obtain background audio text.

[0051] S2044: Find the background audio corresponding to the background audio text.

[0052] This step may include calculating the similarity between the background audio text and pre-stored lyrics, identifying lyrics with a similarity greater than a preset threshold as target lyrics, and identifying the audio corresponding to the target lyrics as background audio; or it may include calculating the similarity between the background audio text and pre-stored lyrics, identifying lyrics with the highest similarity as target lyrics, and identifying the audio corresponding to the target lyrics as background audio.

[0053] S2045: Remove background audio from the audio data to obtain the noise-reduced frequency.

[0054] In this step, the audio data may include taking the negative value of the decibel level of the background audio to obtain the audio to be removed, and superimposing the audio data into the audio to be removed to obtain the noise-reduced frequency; or it may include inputting the audio data and background audio into a pre-trained noise reduction model to obtain the noise-reduced frequency output by the noise reduction model.

[0055] As can be seen from the description of the above embodiments, the embodiments of this disclosure determine the video theme-related text by using fused information text and description text, remove the video theme-related text from the audio text to obtain the background sound text, find the background audio related to the background sound text, remove the background audio from the audio data to obtain the noise-reduced frequency, thereby realizing the removal of background audio from the audio data to obtain the noise-reduced frequency, thus completing the noise reduction of the audio.

[0056] In one possible implementation, step S2041 above involves using fused information text and descriptive text to determine video topic-related text, including: step S41A, step S41B, or step S41C.

[0057] S41A: Combine the informational text and descriptive text to obtain text related to the video theme.

[0058] This step may include concatenating the fusion information text and the description text to obtain video theme-related text, or it may include writing the fusion information and the description text into the same file to obtain video theme-related text.

[0059] S41B: Identify similar texts in the merged information text and description text, and determine the similar texts as texts related to the video theme.

[0060] This step may include calculating the similarity between each field in the fused information and each field in the description text, identifying fields with a similarity greater than a preset similarity threshold as similar fields, and concatenating the similar fields to obtain the video theme-related text.

[0061] S41C: Remove duplicates from the fused information text and description text to obtain deduplicated fused information and deduplicated description text. Combine the deduplicated fused information and deduplicated description text to obtain the text related to the video theme.

[0062] This step involves calculating the similarity between each field in the fused information and each field in the descriptive text. Fields with a similarity greater than a preset similarity threshold are identified as similar fields. Similar fields in the fused information text are then removed to obtain deduplicated fused information. Similarly, similar fields in the descriptive text are removed to obtain deduplicated descriptive text. The deduplicated fused information and deduplicated descriptive text are then concatenated to obtain video-theme-related text. Alternatively, the fused information, deduplicated descriptive text, and preset prompts can be input into a large language model to obtain video-theme-related text after the large language model reorganizes the language.

[0063] As can be seen from the description of the above embodiments, the embodiments of this disclosure obtain video theme-related text by directly combining the fused information text and the description text, or by determining similar text in the fused information text and the description text as video theme-related text, or by combining the deduplicated fused information text and the description text to obtain video theme-related text, thereby combining the data extracted by the fusion model with the description text to obtain actual text related to the theme, which is convenient for subsequent background noise removal.

[0064] In one possible implementation, step S205 above, determining video content information based on the noise-reduced frequency and video data, includes:

[0065] S205A1: Input the noise-reduced audio and video data into the audio-video fusion extraction model to obtain the video content information output by the audio-video fusion extraction model.

[0066] The audio and video fusion extraction model used in this step can be the same as that in step S203 above, or it can consist of an audio preprocessing module, an audio text extraction module, a video text extraction module, and a fusion module.

[0067] The audio preprocessing module is used to decode, reduce noise, and convert the input noise-reduced audio, and then input the processed audio into the audio text extraction module. The audio text extraction module performs speech recognition to obtain speech text information. The video text extraction module performs text recognition in video frames to obtain video text information. The fusion module merges the speech text information and the video text information. The fusion process may include deduplication, semantic analysis, sentence reconstruction, etc.

[0068] As can be seen from the description of the above embodiments, the embodiments of this disclosure obtain video content information by inputting noise-reduced audio and video data into an audio-video fusion extraction model, thereby combining audio and video content to obtain accurate video content information.

[0069] In one possible implementation, step S205 above, determining video content information based on the noise-reduced frequency and video data, includes:

[0070] S205B1: Input the noise-reduced frequency into the audio information extraction model to obtain the noise-reduced frequency information.

[0071] In this step, the audio information extraction model can be trained using methods such as Long Short-Term Memory (LSTM) networks or Transformer models. The noise-reduced audio is then used to perform speech recognition on the noise-reduced audio, yielding the noise-reduced audio information.

[0072] S205B2: Input video data into the video information extraction model to obtain video information.

[0073] In this step, the video information extraction model can use optical character recognition (OCR) to extract text information from each frame of the video data to obtain the video information.

[0074] S205B3: Combine the noise-reduced audio information and video information to obtain video content information.

[0075] This step may include splicing the noise-reduced audio information and video information to obtain video content information, or it may include writing the noise-reduced audio information and video information into the same file to obtain video content information.

[0076] As can be seen from the description of the above embodiments, the embodiments of this disclosure obtain video content information by extracting and combining the noise-reduced frequency information and video information respectively, thereby increasing the accuracy of the video content information.

[0077] In one possible implementation, after step S204 above, which involves fusing the information text and description text, removing background noise from the audio data, and obtaining the noise-reduced audio, the method further includes:

[0078] S220: Identify the noise-reduced frequency information corresponding to the noise-reduced frequency.

[0079] This step is similar to step S205B1 above, and will not be repeated here.

[0080] S221: Determine video content information based on the noise-reduced frequency information, fusion information, and descriptive text.

[0081] This step may include inputting the noise-reduced frequency information, fusion information, and descriptive text into the fusion module to obtain the video content information output by the fusion module. The fusion module is similar to the fusion module in step S205A1 above, and will not be described again here. Alternatively, it may include deduplicating and recombining the noise-reduced frequency information, fusion information, and descriptive text to obtain the video content information.

[0082] As can be seen from the description of the above embodiments, the embodiments of this disclosure extract noise-reduced frequency information, and use the noise-reduced frequency information, fusion information and descriptive text to jointly determine video content information, thereby obtaining more accurate video content information.

[0083] In one possible implementation, after determining the video content information based on the noise-reduced frequency and video data in step S205, the method further includes:

[0084] S230: Establish the correspondence between video content information and the video to be identified.

[0085] This step may include writing video content information and the identifier of the video to be identified into a table or file, thereby establishing a correspondence between video content information and the video to be identified.

[0086] The identifier of the video to be identified may include the video's number, the video's title, etc.

[0087] As can be seen from the description of the above embodiments, the embodiments of this disclosure establish a correspondence between video content information and the video to be identified, which facilitates the subsequent retrieval of the corresponding video through the video content information.

[0088] In one possible implementation, after establishing the correspondence between video content information and the video to be identified in step S230 above, the method further includes:

[0089] S231: Receive prompt word sent by the terminal device.

[0090] In this step, the prompt can be entered by the user into the terminal device, or it can be generated by the terminal device based on the user's input fields. The prompt can be received via messages, data packets, or other means.

[0091] S232: Input the prompt words and target video content information into the video search model to obtain the search results output by the video search model, wherein the target video content information can be any video content information.

[0092] In this step, the video search model can be pre-trained by staff, a large language model, or a BERT model, etc.

[0093] S233: If the search result is yes, then output the video to be identified corresponding to the target video content information.

[0094] In this step, the video to be identified is output, which may include sending the video to be identified to the terminal device, or sending the webpage link of the video to be identified to the terminal device.

[0095] Figure 3 This is a schematic diagram of the structure of the video content recognition device provided in an embodiment of this application. Figure 3 As shown, the video content recognition device 300 includes: a data acquisition module 301, a volume judgment module 302, a text acquisition module 303, a background noise removal module 304, and a content determination module 305.

[0096] The data acquisition module 301 is used to acquire the video to be identified, the comment data of the video to be identified, and the descriptive text, wherein the video to be identified consists of video data and audio data;

[0097] The volume judgment module 302 is used to input comment data into the semantic analysis model to obtain the volume judgment result of the background music;

[0098] The text acquisition module 303 is used to input video data and audio data into the audio-video fusion extraction model if the volume judgment result is that the volume is large, so as to obtain the fusion information text output by the audio-video fusion extraction model.

[0099] Background noise removal module 304 is used to remove background noise from audio data by merging information text and description text to obtain noise-reduced audio.

[0100] The content determination module 305 is used to determine video content information based on the noise-reduced audio and video data.

[0101] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0102] In one possible implementation, the background noise removal module 304 is specifically used to determine the video theme-related text by using the fused information text and description text; determine the audio text based on the audio data; remove the video theme-related text from the audio text to obtain the background noise text; find the background audio corresponding to the background noise text; and remove the background audio from the audio data to obtain the noise-reduced audio.

[0103] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0104] In one possible implementation, the background noise removal module 304 is specifically used to combine the fused information text and the descriptive text to obtain video theme-related text; or, to determine similar text in the fused information text and the descriptive text, and to determine the similar text as video theme-related text; or, to deduplicate the fused information text and the descriptive text to obtain deduplicated fused information and deduplicated descriptive text; and to combine the deduplicated fused information and deduplicated descriptive text to obtain video theme-related text.

[0105] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0106] In one possible implementation, the content determination module 305 is specifically used to input the noise-reduced audio and video data into the audio-video fusion extraction model to obtain the video content information output by the audio-video fusion extraction model.

[0107] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0108] In one possible implementation, the video content recognition device 300 further includes an information determination module 306.

[0109] The information determination module 306 is used to identify the noise-reduced frequency information corresponding to the noise-reduced frequency; and to determine the video content information based on the noise-reduced frequency information, the fusion information and the descriptive text.

[0110] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0111] In one possible implementation, the video content recognition device 300 further includes a relationship establishment module 307.

[0112] The relationship establishment module 307 is specifically used to establish the correspondence between video content information and the video to be identified.

[0113] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0114] In one possible implementation, the video content recognition device 300 further includes a video output module 308.

[0115] The video output module 308 is used to receive prompt words sent by the terminal device; input the prompt words and target video content information into the video search model to obtain the search results output by the video search model, wherein the target video content information is any video content information; if the search result is correct, the video to be identified corresponding to the target video content information is output.

[0116] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0117] To implement the above embodiments, this application also provides an electronic device.

[0118] refer to Figure 4The diagram illustrates a structural schematic of an electronic device 400 suitable for implementing embodiments of this application. The electronic device 400 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0119] like Figure 4 As shown, the electronic device 400 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 401 and a memory 402 communicatively connected to the processor. The processor can perform various appropriate actions and processes based on programs stored in the memory 402, computer-executed instructions, or programs loaded from storage device 408 into random access memory (RAM) 403, to implement the video content recognition method in any of the above embodiments. The memory may be a read-only memory (ROM). The RAM 403 also stores various programs and data required for the operation of the electronic device 400. The processing device 401, the memory 402, and the RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0120] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0121] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from memory 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this application.

[0122] It should be noted that the computer-readable storage medium described above in this application can be a computer-readable signal medium, a computer storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0123] The aforementioned computer-readable storage medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0124] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0125] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0127] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the units do not necessarily limit the module itself; for example, the content determination module can also be described as a "video content information determination module".

[0128] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0129] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the technical solution of the video content recognition method in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the video content recognition method. Please refer to the implementation principle and beneficial effects of the video content recognition method. It will not be repeated here.

[0130] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0131] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the technical solution of the video content recognition method in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the video content recognition method, and can be found in the implementation principle and beneficial effects of the video content recognition method, which will not be repeated here.

[0132] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

[0133] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0134] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method of video content recognition, characterized by, include: Acquire the video to be identified, the comment data of the video to be identified, and the descriptive text, wherein the video to be identified consists of video data and audio data; The comment data is input into a semantic analysis model to obtain the result of judging the volume of the background music; If the volume judgment result is that the volume is large, then the video data and audio data are input into the audio-video fusion extraction model to obtain the fusion information text output by the audio-video fusion extraction model; Using the fused information text and the description text, background noise is removed from the audio data to obtain the noise-reduced audio. Based on the noise-reduced frequency and the video data, determine the video content information.

2. The method of claim 1, wherein, The step of removing background noise from the audio data using the fused information text and the description text to obtain the noise-reduced audio includes: Using the fused information text and the descriptive text, determine the text related to the video theme; Based on the audio data, determine the audio text; Remove the video-related text from the audio text to obtain the background audio text; Find the background audio corresponding to the background audio text; The background audio is removed from the audio data to obtain the noise-reduced audio.

3. The method of claim 2, wherein, The step of determining video topic-related text using the fused information text and the descriptive text includes: The fused information text and descriptive text are combined to obtain video-themed text; or, Identify similar text in the fused information text and the description text, and determine the similar text as text related to the video theme; or, The fused information text and the description text are deduplicated to obtain deduplicated fused information and deduplicated description text; the deduplicated fused information and the deduplicated description text are combined to obtain video theme-related text.

4. The method of claim 1, wherein, The step of determining video content information based on the noise-reduced frequency and the video data includes: The noise-reduced audio and the video data are input into the audio-video fusion extraction model to obtain the video content information output by the audio-video fusion extraction model.

5. The method according to any one of claims 1 to 4, characterized in that, After using the fused information text and the description text to remove background noise from the audio data and obtain the noise-reduced audio, the method further includes: Identify the noise-reduced frequency information corresponding to the noise-reduced frequency; Based on the noise-reduced frequency information, the fusion information, and the descriptive text, the video content information is determined.

6. The method according to any one of claims 1 to 4, characterized in that, After determining the video content information based on the noise-reduced frequency and the video data, the method further includes: Establish the correspondence between the video content information and the video to be identified.

7. The method according to claim 6, characterized in that, After establishing the correspondence between the video content information and the video to be identified, the method further includes: Receive prompts sent by the terminal device; Input the prompt words and target video content information into the video search model to obtain the search results output by the video search model, wherein the target video content information can be any video content information; If the search result is yes, then the video to be identified corresponding to the target video content information will be output.

8. A video content recognition device, characterized in that, include: The data acquisition module is used to acquire the video to be identified, the comment data of the video to be identified, and the descriptive text, wherein the video to be identified consists of video data and audio data; The volume judgment module is used to input the comment data into the semantic analysis model to obtain the volume judgment result of the background music; The text acquisition module is used to input the video data and audio data into the audio-video fusion extraction model if the volume judgment result is that the volume is large, so as to obtain the fusion information text output by the audio-video fusion extraction model. The background noise removal module is used to remove background noise from the audio data using the fused information text and the description text, to obtain noise-reduced audio. The content determination module is used to determine video content information based on the noise-reduced frequency and the video data.

9. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the video content recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the video content recognition method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the video content recognition method as described in any one of claims 1 to 7.