Multimedia advertisement identification method and device
By analyzing and feature extraction of multimedia data, combining language model and advertising library search, the problem of inaccurate advertising recognition in multimedia is solved, and efficient and accurate advertising recognition and monitoring is achieved.
Patent Information
- Application Number
- CN202510159986.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art identification of advertisements in multimedia is not accurate enough, and it is easy to miss advertisements or misidentify non-advertising content as advertisements.
By analyzing the multimedia data, text data is extracted, segmented and encapsulated, and converted into feature vectors. Then, the ads with similarity exceeding the threshold are retrieved in the preset ad library, and the pending text data is extracted in combination with the language model, and finally the ad text is determined based on the similarity.
It improves the accuracy of advertising recognition in multimedia, reduces misidentification, and realizes rapid and effective identification and monitoring of advertisements.
Smart Images

Figure CN120067346A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of advertisement recognition, and particularly relates to a multimedia advertisement recognition method and device. Background Art
[0002] With the rapid development of information technology and the continuous change of the multimedia environment, in order to monitor and manage the advertisement information in multimedia, the advertisement recognition technology for multimedia such as television and radio stations as well as videos has made remarkable development and progress. However, the existing technology for advertisement recognition has insufficient recognition accuracy, and it is easy to miss the advertisement recognition or mis-recognize the multimedia content as advertisement content.
[0003] Therefore, how to improve the accuracy of advertisement recognition in multimedia is a technical problem to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of the present invention is to solve the technical problems of inaccurate and inefficient recognition of advertisements in multimedia in the prior art.
[0005] To achieve the above technical purpose, on the one hand, the present invention provides a multimedia advertisement recognition method, which includes: After parsing the multimedia data, extract a plurality of text data, and the multimedia data is specifically one or a combination of text, pictures, audio and video; Segment the text data and encapsulate it in a preset format to obtain a plurality of groups of text data, and convert each group of text data into a corresponding group feature vector; Retrieve the advertisements in the preset advertisement library whose similarity with each group feature vector exceeds the first threshold, and then form a temporary list in a preset format with the advertisement names and advertisement themes of the advertisements; Extract the text data to be determined from the plurality of groups of text data through a language model in combination with the temporary list, and then encapsulate the text data to be determined in a preset format to obtain a first list; Determine the advertisement text based on the first list and the preset advertisement library.
[0006] Further, when the multimedia data is specifically video or audio, the method further includes extracting keyword groups based on the text data for advertisement recognition.
[0007] Further, the determining the advertisement text based on the first list and the preset advertisement library specifically includes: Merge the text data to be determined with the same advertisement name in the first list and perform vectorization processing to obtain a second list; Compare the data in the second list with the non-advertising database, and delete the data whose similarity exceeds the second threshold to obtain a third list; Determine the advertising text based on the third list and the preset advertising library.
[0008] Further, the determining the advertising text based on the third list and the preset advertising library specifically includes: Determine the data in the third list whose similarity with the data in the preset advertising library exceeds the first similarity threshold as exact advertisements; Determine the data in the third list whose similarity with the data in the preset advertising library is less than the first similarity threshold but greater than the second similarity threshold as pending advertising data; Determine the data in the third list whose similarity with the data in the preset advertising library is less than the second similarity threshold as non-advertising data.
[0009] Further, the merging of the pending text data with the same advertising name in the first list specifically includes: Delete the punctuation marks from all the data in the first list, and then segment the advertising name, advertising theme or advertising slogan in each pending text data respectively to obtain the first phrase and the second phrase corresponding to each pending text data. The first phrase is the phrase after segmenting the advertising name, and the second phrase is the phrase after segmenting the advertising theme or advertising slogan; Determine the intersection words, union words, and the ratio of the number of intersection words to the number of union words between every two of the first phrases, and determine the intersection words, union words, and the ratio of the number of intersection words to the number of union words between every two of the second phrases; Merge the pending text data with exactly the same first phrase, and / or merge the pending text data corresponding to the ratio of the number of intersection words to the number of union words between every two of the first phrases being greater than the specified threshold, and / or if the intersection words between every two of the first phrases are greater than zero and the ratio of the number of intersection words to the number of union words between the corresponding second phrases of the corresponding pending text data is greater than the specified threshold, merge the corresponding two groups of pending text data.
[0010] Further, the determining the advertising text based on the first list and the preset advertising library specifically includes: Divide each pending text data in the first list into corresponding sentences, and save all the sentences and their time information to the search server; Select keywords from each advertising sample in the preset advertising library to form a keyword group; Fuzzily search each keyword in the keyword group in a search server to obtain corresponding search records and their first-time information, where the first-time information is specifically the timestamp information of the search records; Sort the search records according to the first-time information to obtain a sequence list; Initialize a sliding window, and sequentially perform advertisement matching on the search records in the sequence list with the sliding window.
[0011] Further, the sequentially performing advertisement matching on the search records in the sequence list with the sliding window specifically includes: When the sliding window slides to the current search record in the sequence list, obtain the keyword corresponding to the current search record, and store the keyword and its timestamp in a keyword temporary storage list; Judge whether the ratio of the number of types of different keywords in the keyword temporary storage list to the number of keywords in the keyword group reaches a preset ratio. If so, determine that the multiple keywords currently obtained by the sliding window are advertisements, initialize the sliding window, and then continue to slide the sliding window. If not, continue to slide the sliding window.
[0012] Further, the text data includes attribute information and time information. The attribute information specifically includes a serial number, text content, and a speaking role. Among them, if there is no time information in the multimedia data corresponding to the text data, only the attribute information is added without adding time information when extracting the text data.
[0013] Further, the splitting of the text data specifically means dividing two text data with a time interval less than a preset time interval between adjacent text data into the same group. The time start point of the time information of the group of text data is the smallest time start point in the group of text data, and the time end point of the time information of the group of text data is the largest time end point in the group of text data.
[0014] Further, after extracting multiple text data, the method further includes filtering out the modal particles in all the text data.
[0015] On the other hand, the present invention also provides a multimedia advertisement recognition device, and the device includes: A parsing module, configured to parse multimedia data and extract multiple text data therefrom. The multimedia data is specifically one or more combinations of text, pictures, audio, and video; A partitioning module, configured to partition the text data, package it according to a preset format to obtain multiple groups of text data, and convert each group of text data into a corresponding group feature vector; A retrieval module, configured to retrieve advertisements in a preset advertisement library that have a similarity exceeding a first threshold with each of the group feature vectors, and then form a temporary list in a preset format with the advertisement names and advertisement themes of these advertisements; An extraction module, configured to extract pending text data from multiple groups of text data through a language model in combination with the temporary list, and then encapsulate the pending text data in a preset format to obtain a first list; A determination module, configured to determine advertisement text based on the first list and the preset advertisement library.
[0016] A multimedia advertisement recognition method and device provided by the present invention, compared with the prior art, first parses multimedia data and then extracts multiple text data, where the multimedia data is specifically one or a combination of text, pictures, audio, and video; divides the text data and encapsulates it in a preset format to obtain multiple groups of text data, and converts each group of text data into a corresponding group feature vector; retrieves advertisements in a preset advertisement library that have a similarity exceeding a first threshold with each of the group feature vectors, and then forms a temporary list in a preset format with the advertisement names and advertisement themes of these advertisements; extracts pending text data from multiple groups of text data through a language model in combination with the temporary list, and then encapsulates the pending text data in a preset format to obtain a first list; determines advertisement text based on the first list and the preset advertisement library. It can quickly and effectively identify advertisements in multimedia, so as to monitor the advertisements. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 The figure shows a schematic flow chart of the multimedia advertisement recognition method provided by the embodiment of this specification; Figure 2 The figure shows a schematic structural diagram of the multimedia advertisement recognition device provided by the embodiment of this specification. Detailed Embodiments
[0019] To enable those of ordinary skill in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0020] As Figure 1 shown in the flowchart of the multimedia advertisement recognition method provided in the embodiments of this specification, although this specification provides the method operation steps or device structures shown in the following embodiments or drawings, based on routine or without creative efforts, more or fewer operation steps or module units may be included in the method or device. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or drawings of this specification. When the described method or module structure is applied to an actual device, server or terminal product, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, even including an environment of distributed processing and server clusters).
[0021] The multimedia advertisement recognition method provided in the embodiments of this specification can be applied to terminal devices such as clients and servers. As Figure 1 shown, the method specifically includes the following steps: Step S101: After parsing the multimedia data, extract multiple text data. The multimedia data is specifically one or a combination of text, pictures, audio, and video.
[0022] The text data includes attribute information and time information. The attribute information specifically includes a serial number, text content, and speaking role. Among them, if there is no time information in the multimedia data corresponding to the text data, only the attribute information is added without adding time information when extracting the text data. After extracting multiple text data, the method further includes filtering out the modal particles in all the text data.
[0023] Specifically, the multimedia data includes one or more combinations of text, pictures, audio, and video. For the speech in audio or video, each sentence needs to be extracted. For the images in the video, after extracting the key frames of the video by the residual method, text extraction is performed on the key frames. At the same time, attribute information and time information are added to the extracted text data. The time information of the text data extracted from the speech data is the start time point and the end time point of the corresponding sentence. The time information of the text data extracted from the video data is the time point of the corresponding key frame.
[0024] In addition, the text data can specifically be in the JSON data structure format, including all the text and its attribute information and time information. After extracting the text data, various modal particles also need to be filtered to clean up and reduce redundant information in the text.
[0025] Step S102: Split the text data and encapsulate it in a preset format to obtain multiple groups of text data, and convert each group of text data into a corresponding group feature vector.
[0026] Specifically, splitting the text data means dividing two text data with a time interval less than a preset time interval between adjacent text data into the same group. The start time of the time information of the group of text data is the smallest start time in this group of text data, and the end time of the time information of the group of text data is the largest end time in this group of text data.
[0027] Specifically, after extracting the text data, adjacent text data with a time interval less than the preset time interval are merged to split out multiple groups of text data. For example, subtract the end time of the first text data from the start time of the second text data. If the result is less than or equal to 30 seconds, then splice the text contents of the two text data with a comma, and at the same time retain the start time of the first text data and the end time of the second text data as the time marker of the group of text data. This merging process will continue until the time intervals of all adjacent text data exceed 30 seconds.
[0028] Step S103: Retrieve the advertisements in the preset advertisement library whose similarity with each group feature vector exceeds the first threshold, and then form a temporary list in a preset format with the advertisement names and advertisement themes of these advertisements.
[0029] Specifically, for each group of text data after merging, convert its text content into a feature vector, and use the cosine similarity algorithm to perform similarity retrieval in the preset advertisement library. Retrieve all advertisements with a similarity reaching or exceeding 80%, and record the name (ad_name) and theme (ad_theme) of each advertisement. Organize and summarize the retrieval results in the form of a JSON list to obtain a temporary list.
[0030] Step S104: Extract the to-be-determined text data from multiple groups of text data through a language model in combination with the temporary list, and then encapsulate the to-be-determined text data in a preset format to obtain a first list.
[0031] Specifically, the language model is corrected through the advertisement name (ad_name) and theme (ad_theme), and then the to-be-determined text data is extracted from multiple groups of text data through the language model. That is, after the language model is corrected and enhanced through the temporary list, the to-be-determined text data is extracted from multiple groups of text data. At the same time, the corresponding attribute information and time information are supplemented for the to-be-determined text data to form a first list, and the first list also includes the advertisement name and advertisement theme.
[0032] Step S105: Determine the advertisement text based on the first list and a preset advertisement library.
[0033] In the embodiment of the present application, the determining the advertisement text based on the first list and a preset advertisement library specifically includes: Merge the to-be-determined text data with the same advertisement name in the first list and perform vectorization processing to obtain a second list; Compare the similarity of each data in the second list with a non-advertisement database, and delete the data whose similarity exceeds a second threshold to obtain a third list; Determine the advertisement text according to the similarity of each data in the third list with the preset advertisement library.
[0034] Specifically, after obtaining the first list, data cleaning is also performed on it. The to-be-determined text data with the same advertisement name and a time interval within three minutes is merged and then vectorized to obtain a second list. Then, each data in the second list, that is, vector data, is compared for similarity in the non-advertisement database through the cosine similarity algorithm. When the similarity exceeds a second threshold, for example, 95%, it means that these vector data are non-advertisement data, and they are deleted to obtain a third list. Then, the cosine similarity of each vector data in the third list is compared in the preset advertisement library to determine the advertisement text.
[0035] More specifically, the merging of the to-be-determined text data with the same advertisement name in the first list specifically includes: Delete the punctuation marks from all data in the first list, and then perform word segmentation on the advertisement name, advertisement theme, or advertisement slogan in each to-be-determined text data to obtain a first word group and a second word group corresponding to each to-be-determined text data. The first word group is the word group after word segmentation of the advertisement name, and the second word group is the word group after word segmentation of the advertisement theme or advertisement slogan. Determine the intersection words, union words, and the ratio of the number of intersection words to the number of union words between the first word groups, and determine the intersection words, union words, and the ratio of the number of intersection words to the number of union words between the second word groups; The pending text data that are exactly the same as the first phrases are merged, and or, the pending text data corresponding to the ratio of the number of intersection words and the number of union words between the first phrases is greater than a specified threshold are merged, and or, if the number of intersection words between the first phrases is greater than zero and the ratio of the number of intersection words and the number of union words between the second phrases corresponding to the corresponding pending text data is greater than a specified threshold, the corresponding two groups of pending text data are merged.
[0036] In practical applications, the solution of this application chooses to perform the merging operation through a language model. Since the input and output of the language model are limited by the number of tokens, it is impossible to feed all the data into the model at once, so it will be cut into multiple segments in the front, and multiple large language model inferences will be performed, and multiple results will be returned. Because this strategy will lead to duplication of the results obtained by language model inference, and there is a problem of incomplete description in the name and subject, so the solution of this application also includes first deleting punctuation marks from each data in the first list and then segmenting the advertisement name and advertisement subject or advertisement slogan separately, wherein the deletion of punctuation marks can also be done after the segmentation process; then calculate the intersection vocabulary, union vocabulary and the ratio of the number of intersection vocabulary and the number of union vocabulary of the advertisement name after segmentation, the intersection vocabulary, union vocabulary and the ratio of the number of intersection vocabulary and the number of union vocabulary of the advertisement subject or advertisement slogan after segmentation; then merge the pending text data in the following three situations: 1. If the first phrases are the same, the corresponding pending text data are merged, and the advertising themes of the corresponding pending text data are merged and spliced; 2. If the intersection-and-union ratio between the first phrases is greater than a specified threshold (e.g., 0.7, which can be flexibly adjusted by a person skilled in the art according to actual conditions), the corresponding pending text data are merged, and the merged advertisement name is a relatively complete (longer) advertisement name, and the merged theme is also a relatively complete (longer) advertisement theme; 3. If the number of intersection words between the first phrases is > 0, and the intersection ratio between the second phrases is > the specified threshold (such as 0.7, which can be adjusted), the corresponding two pending text data will be merged, and the merged advertising name will be a more complete (longer) advertising name, and the merged theme will also be a more complete advertising theme.
[0037] Example: To-be-determined text data 1: Advertisement name: Nongfu Spring, advertisement theme: We are the porters of nature, start time: 10:10:10, end time: 10:10:20; To-be-determined text data 2: Advertisement name: Nongfu Spring Mineral Water, advertisement theme: We are not the producers of water, start time: 10:10:30, end time: 10:10:40; After processing the two to-be-determined text data according to the above word segmentation, the first phrase corresponding to the to-be-determined text data 1 includes "Nongfu Spring", and the corresponding second phrase includes "We", "are", "nature", "of", "porters"; the first phrase corresponding to the to-be-determined text data 2 includes "Nongfu Spring", "Mineral Water", and the corresponding second phrase includes "We", "are not", "water", "of", "producers"; the intersection word among the first phrases is "Nongfu Spring", the union words among the first phrases are "Nongfu Spring", "Mineral Water"; the intersection words among the second phrases are "We", "of", and the union words among the second phrases are "We", "are", "nature", "of", "porters", "are not", "water", "producers".
[0038] And in the above three cases, during merging, advertisements without English in the advertisement name are preferentially merged, and advertisements with more characters in the advertisement name are preferentially merged, and the start and end times of the merged to-be-determined text data are updated.
[0039] In some embodiments, determining the advertisement text based on the third list and the preset advertisement library specifically includes: Determining the data in the third list with a similarity to the data in the preset advertisement library exceeding the first similarity threshold as the exact advertisement; Determining the data in the third list with a similarity to the data in the preset advertisement library less than the first similarity threshold but greater than the second similarity threshold as the to-be-determined advertisement data; Determining the data in the third list with a similarity to the data in the preset advertisement library less than the second similarity threshold as the non-advertisement data.
[0040] Specifically, the data in the third list with a similarity to the data in the preset advertisement library exceeding the first similarity threshold, such as 95% or 90%, is determined as the exact advertisement. If the similarity threshold is less than the first similarity threshold and greater than the second similarity threshold, for example, greater than 65% and less than 95%, the corresponding data is determined as the to-be-determined advertisement data, that is, the suspected advertisement. If the similarity threshold is less than the second similarity threshold, such as less than 65%, the corresponding data is determined as the non-advertisement.
[0041] In some embodiments, when the multimedia data is specifically video or audio, keyword groups are also extracted based on the text data for advertisement recognition, which specifically includes: Divide the corresponding statements from each pending text data in the first list, and save all the statements and their time information to the search server; Select keywords from each advertisement sample in the preset advertisement library to form keyword groups; Perform fuzzy searches on each keyword in the keyword group in the search server to obtain the corresponding search records and their first time information, where the first time information is specifically the timestamp information of the search records; Sort the search records according to the first time information to obtain a sequence list; Initialize a sliding window, and perform advertisement matching on the search records in the sequence list in turn with the sliding window.
[0042] Specifically, that is, in the solution of the present application, when finally matching advertisements, it can also be to divide the corresponding statements and their time information from each pending text data in the first list. This time information is also the start and end time of the statement. Save each statement and its time information as a separate record in the search server, then extract the advertisement samples, extract the keywords in the advertisement samples to form keyword groups. Optionally, the keywords in the keyword group are separated by "|", and count the number of keywords in the keyword group. Then perform fuzzy searches on each keyword in the keyword group in the search server to obtain the corresponding search records and their first time information. This first time information is also the timestamp information. Then sort these search records and their first time information according to the timestamp to obtain a sequence list, and finally perform advertisement matching in the sequence list through the sliding window.
[0043] The step of performing advertisement matching on the search records in the sequence list in turn with the sliding window specifically includes: When the sliding window slides to the current search record in the sequence list, obtain the keyword corresponding to the current search record, and store the keyword and its timestamp in the keyword temporary storage list; Judge whether the ratio of the number of types of different keywords in the keyword temporary storage list to the number of keywords in the keyword group reaches a preset ratio. If so, determine that the multiple keywords currently obtained by the sliding window are advertisements, initialize the sliding window, and then continue to slide the sliding window. If not, continue to slide the sliding window.
[0044] Specifically, when using a sliding window to perform sliding matching in a sequence list, first initialize the sliding window, including setting the start and end times of the sliding window to 0 and clearing the keyword temporary storage list. Then, slide the sliding window in the sequence list for matching. When the sliding window slides to the current search record, first determine whether the time interval between each keyword in the keyword temporary storage list is less than the first preset time interval. If so, update the start and end times of the sliding window according to the start and end times of each keyword. If not, determine whether the ratio of the number of different keyword types in the keyword temporary storage list to the number of keywords in the keyword group reaches the preset ratio. If so, determine that multiple keywords in the keyword temporary storage list are advertisements, and the advertisement time is the start time of the first keyword to the end time of the last keyword. Then, initialize the window, add the current keyword to the keyword list, and update the start and end times of this record to the start and end times of the window. Then, continue to slide the sliding window. If not, continue to slide the sliding window to match the next search record.
[0045] Based on the above multimedia advertisement recognition method, one or more embodiments of this specification also provide a multimedia advertisement recognition platform and terminal. The platform or terminal may include devices, software, modules, plugins, servers, clients, etc. that use the method described in the embodiments of this specification and combine necessary implementation hardware devices. Based on the same innovative concept, the systems in one or more embodiments provided in the embodiments of this specification are as described in the following embodiments. Since the implementation schemes for the systems to solve problems are similar to the method, the implementation of the specific systems in the embodiments of this specification may refer to the implementation of the foregoing method, and repeated parts will not be elaborated. The term "unit" or "module" used hereinafter may be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated.
[0046] Specifically, Figure 2 is a schematic module structure diagram of an embodiment of the multimedia advertisement recognition device provided in this specification. As Figure 2 shown, the multimedia advertisement recognition device provided in this specification includes: A parsing module 201, configured to extract multiple text data after parsing the multimedia data, where the multimedia data is specifically one or more combinations of text, pictures, audio, and video; A partitioning module 202, configured to split the text data, encapsulate it in a preset format to obtain multiple groups of text data, and convert each group of text data into a corresponding group feature vector; A retrieval module 203, configured to retrieve advertisements in a preset advertisement library whose similarity to each group feature vector exceeds a first threshold, and then form a temporary list in a preset format with the advertisement names and advertisement themes of these advertisements; An extraction module 204, configured to extract to-be-determined text data from multiple groups of text data through a language model in combination with the temporary list, and then encapsulate the to-be-determined text data in a preset format to obtain a first list; A determination module 205, configured to determine advertisement texts based on the first list and a preset advertisement library.
[0047] It should be noted that the above system may further include other implementation manners according to the description of the corresponding method embodiments. The specific implementation manners may refer to the description of the corresponding method embodiments above, and will not be elaborated here one by one.
[0048] An embodiment of the present application further provides an electronic device, including: A processor; A memory for storing executable instructions of the processor; The processor is configured to execute the method provided in the above embodiment.
[0049] The electronic device provided in the embodiment of the present application, by storing the executable instructions of the processor in the memory, when the processor executes the executable instructions, can first parse multimedia data and then extract multiple text data, where the multimedia data is specifically one or a combination of text, pictures, audio, and video; split the text data and encapsulate it in a preset format to obtain multiple groups of text data, and convert each group of text data into a corresponding group feature vector; retrieve advertisements in a preset advertisement library whose similarity with each group feature vector exceeds a first threshold, and then form a temporary list in a preset format with the advertisement names and advertisement themes of the advertisements; extract to-be-determined text data from multiple groups of text data through a language model in combination with the temporary list, and then encapsulate the to-be-determined text data in a preset format to obtain a first list; determine advertisement texts based on the first list and a preset advertisement library. It can quickly and effectively identify advertisements in multimedia, so as to monitor the advertisements.
[0050] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0051] The method or device described in the above embodiments provided in this specification can implement business logic through a computer program and record it on a storage medium. The storage medium can be read and executed by a computer to achieve the effects of the solutions described in the embodiments of this specification, such as: After parsing the multimedia data, multiple text data are extracted. The multimedia data is specifically one or a combination of text, pictures, audio, and video; The text data is segmented and encapsulated in a preset format to obtain multiple groups of text data, and each group of text data is converted into a corresponding group feature vector; Advertisements with a similarity exceeding a first threshold to each of the group feature vectors are retrieved from a preset advertisement library, and then the advertisement names and advertisement themes of these advertisements are formed into a temporary list in a preset format; Pending text data is extracted from multiple groups of text data through a language model in combination with the temporary list, and then the pending text data is encapsulated in a preset format to obtain a first list; Based on the first list and the preset advertisement library, advertisement text is determined.
[0052] The storage medium may include a physical device for storing information. Usually, information is digitized and then stored in a medium using electrical, magnetic, or optical means. The storage medium may include: devices that store information using electrical energy, such as various memories, such as RAM, ROM, etc.; devices that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, bubble memories, USB flash drives; devices that store information using optical means, such as CDs or DVDs. Of course, there are also other types of readable storage media, such as quantum memories, graphene memories, and so on.
[0053] The embodiments of this specification are not limited to those that must conform to industry communication standards, standard computer resource data update and data storage rules, or the situations described in one or more embodiments of this specification. Certain industry standards or implementation schemes slightly modified based on the implementation described in a custom manner or embodiment can also achieve the same, equivalent, or similar, or predictable implementation effects after deformation as the above embodiments. The embodiments obtained by applying these modified or deformed data acquisition, storage, judgment, processing methods, etc. still fall within the scope of the optional implementation schemes of the embodiments of this specification.
[0054] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0055] The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or plugins can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0056] These computer program instructions can also be loaded onto a computer or other programmable resource data update device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.
[0057] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0058] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A multimedia advertisement recognition method, characterized in that: The method comprises: Extracting a plurality of text data after parsing the multimedia data, wherein the multimedia data is specifically one or more combinations of text, picture, audio and video; Segmenting the text data and packaging them according to a preset format to obtain a plurality of groups of text data, and converting each group of text data into a corresponding group feature vector; Retrieving advertisements whose similarity with each of the groups of feature vectors exceeds a first threshold value from a preset advertisement library, and then composing the advertisement names and advertisement themes of the advertisements into a temporary list in a preset format; Extracting pending text data from multiple groups of text data by using a language model and combining the temporary list, and then packaging the pending text data in a preset format to obtain a first list; The advertisement text is determined based on the first list and a preset advertisement library.
2. The multimedia advertisement recognition method according to claim 1, characterized in that: The method further includes extracting a keyword group based on the text data to perform advertisement identification when the multimedia data is specifically video or audio.
3. The multimedia advertisement recognition method according to claim 1, characterized in that: The step of determining the advertisement text based on the first list and the preset advertisement library specifically includes: Merging the pending text data with the same advertisement name in the first list and performing vectorization processing to obtain a second list; Compare the similarity of each data in the second list with the non-advertisement database, and delete the data whose similarity exceeds the second threshold to obtain a third list; The advertisement text is determined based on the third list and a preset advertisement library.
4. The multimedia advertisement recognition method according to claim 3, characterized in that: The determining of the advertisement text based on the third list and the preset advertisement library specifically includes: Determine the data in the third list whose similarity with the data in the preset advertisement library exceeds a first similarity threshold as the correct advertisement; Determine the data in the third list whose similarity with the data in the preset advertisement library is less than the first similarity threshold but greater than the second similarity threshold as pending advertisement data; The data in the third list whose similarity with the data in the preset advertisement library is less than the second similarity threshold is determined as non-advertisement data.
5. The multimedia advertisement recognition method according to claim 3, characterized in that: The step of merging the pending text data with the same advertisement name in the first list specifically includes: After deleting punctuation marks from all data in the first list, segmenting the advertisement name, advertisement theme or advertisement slogan in each pending text data to obtain a first phrase and a second phrase corresponding to each pending text data, wherein the first phrase is a phrase after segmentation of the advertisement name, and the second phrase is a phrase after segmentation of the advertisement theme or advertisement slogan; Determine the intersection words, union words, and the ratio of the number of intersection words to the number of union words between the first word groups, and determine the intersection words, union words, and the ratio of the number of intersection words to the number of union words between the second word groups; The pending text data that are exactly the same as the first phrases are merged, and / or the pending text data corresponding to the ratio of the number of intersection words to the number of union words between the first phrases is greater than a specified threshold value are merged, and / or, if the number of intersection words between the first phrases is greater than zero and the ratio of the number of intersection words to the number of union words between the second phrases corresponding to the corresponding pending text data is greater than a specified threshold value, the corresponding two groups of pending text data are merged.
6. The multimedia advertisement recognition method according to claim 2, characterized in that: The step of determining the advertisement text based on the first list and the preset advertisement library specifically includes: Divide each pending text data in the first list into corresponding sentences, and save all the sentences and their time information to the search server; Select keywords from each advertisement sample in the preset advertisement library to form a keyword group; Performing fuzzy search on each keyword in the keyword group in the search server to obtain corresponding search records and first time information thereof, wherein the first time information is specifically timestamp information of the search records; Sorting the search records according to the first time information to obtain a sequence list; Initialize a sliding window, and use the sliding window to match advertisements to the search records in sequence.
7. The multimedia advertisement recognition method according to claim 6, characterized in that: The step of sequentially matching the search records with advertisements using the sliding window in the sequence table specifically includes: When the sliding window slides to the current search record in the sequence table, a keyword corresponding to the current search record is obtained, and the keyword and its timestamp are stored in a keyword temporary storage list; Determine whether the ratio of the number of different keyword types in the keyword temporary storage list to the number of keywords in the keyword group reaches a preset ratio. If so, determine that the multiple keywords currently acquired by the sliding window are advertisements, and continue to slide the sliding window after initializing the sliding window. If not, continue to slide the sliding window.
8. The multimedia advertisement recognition method according to claim 1, characterized in that: The text data includes attribute information and time information, and the attribute information specifically includes a sequence number, text content, and a speaking role. If the multimedia data corresponding to the text data does not contain time information, only attribute information is added when the text data is extracted without adding time information.
9. The multimedia advertisement recognition method according to claim 1, characterized in that: The text data segmentation is specifically to divide two text data whose time interval between adjacent text data is less than a preset time interval into the same group, the time starting point of the time information of the group of text data is the smallest time starting point in the group of text data, and the time ending point of the time information of the group of text data is the largest time ending point in the group of text data.
10. A multimedia advertisement recognition device, characterized in that: The device comprises: A parsing module, used to parse the multimedia data and extract multiple text data, wherein the multimedia data is specifically one or more combinations of text, pictures, audio and video; A division module, used for dividing the text data and encapsulating them according to a preset format to obtain a plurality of groups of text data, and converting each group of text data into a corresponding group feature vector; A retrieval module, used to retrieve advertisements whose similarity with each of the groups of feature vectors exceeds a first threshold value in a preset advertisement library, and then compose the advertisement names and advertisement themes of the advertisements into a temporary list in a preset format; An extraction module, used for extracting pending text data from a plurality of groups of text data by using a language model and combining the temporary list, and then encapsulating the pending text data in a preset format to obtain a first list; The determination module is used to determine the advertisement text based on the first list and the preset advertisement library.