Brand word acquisition method and device, electronic equipment and storage medium

By comprehensively utilizing multimodal detection methods that combine image and text information from notes, the problem of low logo detection accuracy was solved, achieving high coverage and high accuracy detection of brand words.

CN117786038BActive Publication Date: 2026-04-28SHUXING TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHUXING TECH (BEIJING) CO LTD
Filing Date
2022-09-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Current technologies for logo detection identify brand terms that are limited in type and have low accuracy, failing to fully cover users' brand intent.

Method used

By combining note images and text information for brand word detection, and utilizing subject recognition, text detection, and brand word fusion, multimodal brand words are obtained, thereby improving detection coverage and accuracy.

Benefits of technology

It achieves comprehensive detection of various types of brand keywords, improves the coverage and accuracy of brand keyword detection, and outputs coordinated target brand keywords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117786038B_ABST
    Figure CN117786038B_ABST
Patent Text Reader

Abstract

The application provides a brand word acquisition method and device, electronic equipment and a storage medium. The method comprises the following steps: obtaining note information, wherein the note information comprises a note picture and a note text; performing subject identification on the note picture to obtain h categories of h subjects, the h subjects and the h categories being in one-to-one correspondence; performing text detection on the note text to obtain a first brand word set; performing brand word detection on the note picture to obtain at least one second brand word set; and performing fusion on the first brand word set and the at least one second brand word set according to the h categories of the h subjects to obtain z target brand words. The embodiment of the application is beneficial to improving brand word detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for obtaining brand terms. Background Technology

[0002] Currently, the application's sharing community features user-shared notes that users can browse. When users are interested in the content of a note, they can take a screenshot of it. Based on the user's screenshot behavior, the app can obtain the user's brand intent, i.e., which brands the user is interested in. This intelligent real-time conversion from screenshot behavior to brand intent can help merchants build daily sales and promotions.

[0003] However, currently, when a user's screenshot behavior is detected, the brand detection is performed on the image content that the user wants to screenshot through the logo to obtain the brand keywords that the screenshot behavior is interested in.

[0004] However, the brand word types identified by logo detection are limited and the accuracy is low. Summary of the Invention

[0005] This application provides a brand word acquisition method, device, electronic device, and storage medium. It performs brand word detection and fusion by comprehensively combining text and images, and can detect various types of brands. It has a high coverage rate of brand words and high detection accuracy.

[0006] In a first aspect, embodiments of this application provide a method for obtaining brand terms, the method comprising:

[0007] In response to a user's screenshot action, obtain note information corresponding to the screenshot action, wherein the note information includes note images and note text;

[0008] Subject recognition is performed on the note image to obtain h subjects and h categories, with each of the h subjects and h categories corresponding one-to-one;

[0009] Text detection is performed on the notes to obtain a first set of brand keywords;

[0010] Brand word detection is performed on the notes image to obtain at least one second brand word set;

[0011] Based on the h categories of the h subjects, the first brand word set and the at least one second brand word set are merged to obtain z target brand words.

[0012] In conjunction with the first aspect, in one possible implementation, the step of fusing the first brand term set and the at least one second brand term set according to the h categories of the h subjects to obtain z target brand terms includes:

[0013] Determine m first target brand words from the first brand word set and the at least one second brand word set, wherein each first target brand word exists in at least two brand word sets, the first brand word set and the at least one second brand word set;

[0014] Remove the m first target brand words from the first brand word set and the at least one second brand word set to obtain a third brand word set corresponding to the first brand word set, and a fourth brand word set corresponding to each second brand word set;

[0015] Based on the h categories, the third brand word set is filtered to obtain n second target brand words;

[0016] Based on the h categories, brand words are filtered for each fourth brand word set to obtain k third target brand words corresponding to each fourth brand word set;

[0017] The m first target brand words, the n second target brand words, and the k third target brand words corresponding to each fourth brand word set are merged to obtain the z target brand words.

[0018] In conjunction with the first aspect, in one possible implementation, the step of filtering the third brand term set according to the h categories to obtain n second target brand terms includes:

[0019] Obtain x fourth target brand words from the third brand word set, wherein the h categories contain the category corresponding to each fourth target brand word;

[0020] Remove the x fourth target brand terms from the third brand term set to obtain the fifth brand term set;

[0021] A weighted processing operation is performed on each brand word in the fifth brand word set according to its corresponding category to obtain the target confidence level of each brand word in the fifth brand word set;

[0022] Based on the target confidence level of each brand word in the fifth brand word set, y fifth target brand words are obtained;

[0023] The x fourth target brand words and the y fifth target brand words are merged to obtain the n second target brand words.

[0024] In conjunction with the first aspect, in one possible implementation, the step of filtering brand words for each fourth brand word set according to the h categories to obtain k third target brand words corresponding to each fourth brand word set includes:

[0025] Obtain q sixth target brand words from each fourth brand word set, wherein the h categories contain the category corresponding to each sixth target brand word;

[0026] Remove the q sixth target brand words from each fourth brand word set to obtain the sixth brand word set corresponding to each fourth brand word set;

[0027] A weighted processing operation is performed on each brand word in the sixth brand word set according to its corresponding category to obtain the target confidence level of each brand word in the sixth brand word set;

[0028] Based on the target confidence level of each brand word in the sixth brand word set, determine r seventh target brand words corresponding to each fourth brand word set;

[0029] The q sixth target brand words in each fourth brand word set and the r seventh target brand words corresponding to each fourth brand word set are merged to obtain the k third target brand words corresponding to each fourth brand word set.

[0030] In conjunction with the first aspect, in one possible implementation, the weighted processing operation includes:

[0031] Based on the category corresponding to brand word A, determine the penalty coefficient corresponding to brand word A, wherein brand word A is any one of the fifth brand word set or any one of each sixth brand word set;

[0032] Based on the brand word set in which brand word A is located, obtain the preset penalty item corresponding to brand word A;

[0033] The target confidence level of brand term A is determined based on the penalty coefficient corresponding to brand term A, the preset penalty item, and the confidence level of brand term A.

[0034] In conjunction with the first aspect, in one possible implementation, the step of performing text detection on the first note text to obtain the first brand keyword set includes:

[0035] Entity recognition is performed on the note text to obtain a first candidate brand words;

[0036] Map each first-candidate brand term to a main brand term to obtain the brand term corresponding to each first-candidate brand term.

[0037] The brand words corresponding to each first candidate brand word are combined into the first brand word set.

[0038] In conjunction with the first aspect, in one possible implementation, the step of performing brand word detection on the note image to obtain at least one second set of brand words includes:

[0039] Brand logo detection was performed on the note image to obtain b second candidate brand words;

[0040] Map each second candidate brand term to the main brand term to obtain the brand term corresponding to each second candidate brand term.

[0041] The brand words corresponding to each second candidate brand word are combined into at least one set of second brand words.

[0042] In conjunction with the first aspect, in one possible implementation, the step of performing brand word detection on the note image to obtain at least one second set of brand words includes:

[0043] Optical character recognition was performed on the note image to obtain c third candidate brand words;

[0044] For each third candidate brand term, perform brand term mapping to obtain the brand term corresponding to each third candidate brand term;

[0045] The brand words corresponding to each third candidate brand word are combined into at least one set of second brand words.

[0046] In conjunction with the first aspect, in one possible implementation, the step of performing brand word detection on the note image to obtain at least one second set of brand words includes:

[0047] Logo detection was performed on the note image to obtain b second candidate brand words;

[0048] Map each second candidate brand term to the main brand term to obtain the brand term corresponding to each second candidate brand term.

[0049] Optical character recognition was performed on the note image to obtain c third candidate brand words;

[0050] Map each second candidate brand term to a main brand term to obtain the brand term corresponding to each third candidate brand term.

[0051] Each second candidate brand word is combined into a second brand word set, and each third candidate brand word is combined into another second brand word set to obtain the at least one second brand word set.

[0052] In conjunction with the first aspect, in one possible implementation, the optical character recognition (OCR) of the note image yields c third candidate brand words, including:

[0053] Optical character recognition is performed on the note image to obtain text information;

[0054] The text information is matched with the brand terminology database to obtain d fourth candidate brand terms;

[0055] The text information is matched with the h categories to obtain e fifth candidate brand words;

[0056] The d fourth candidate brand words and the e fifth candidate brand words are used as the c third candidate brand words.

[0057] In conjunction with the first aspect, in one possible implementation, the step of performing subject recognition on the note image to obtain h categories of h subjects includes:

[0058] Object detection is performed on the note image to obtain multiple candidate boxes;

[0059] Based on the position information of each candidate box, subject detection is performed on the multiple candidate boxes to obtain h subject boxes, wherein the target selected by each subject box is a subject;

[0060] The subject corresponding to each subject frame is classified to obtain h categories corresponding to each subject.

[0061] In conjunction with the first aspect, in one possible implementation, the note information further includes note categories;

[0062] The note category is used to provide prior information when performing text detection on the note text, so that the brand words in the obtained first brand word set can be matched with the note category;

[0063] The note category is also used to provide prior information when performing brand word detection on the note image, so that the brand words in each second brand word set obtained are matched with the note category;

[0064] The note category is also used to provide prior information when performing subject recognition on the note image, so that each subject obtained, and the category of each subject, matches the note category.

[0065] In conjunction with the first aspect, in one possible implementation, the method further includes:

[0066] Based on the z target brand keywords, brand recommendations are made to the user.

[0067] In conjunction with the first aspect, in one possible implementation, the method further includes:

[0068] Based on the z target brand keywords, a user profile is constructed for the user.

[0069] Secondly, embodiments of this application provide a brand term acquisition device, including:

[0070] The acquisition unit is used to respond to the user's screenshot behavior and acquire note information corresponding to the screenshot behavior, wherein the note information includes note images and note text;

[0071] The processing unit is used to perform subject recognition on the note image to obtain h categories of h subjects, wherein the h subjects and the h categories correspond one-to-one;

[0072] Text detection is performed on the notes to obtain a first set of brand keywords;

[0073] Brand word detection is performed on the notes image to obtain at least one second brand word set;

[0074] Based on the h categories of the h subjects, the first brand word set and the at least one second brand word set are merged to obtain z target brand words.

[0075] Thirdly, embodiments of this application provide an electronic device, a processor, and a memory, wherein the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in the first aspect.

[0076] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that causes a computer to perform the method described in the first aspect.

[0077] Fifthly, embodiments of this application provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, the computer being operable to perform the method as described in the first aspect.

[0078] The above-mentioned solution in this application includes at least the following beneficial effects:

[0079] In this embodiment, firstly, in response to a user's screenshot action, note information corresponding to that screenshot action is obtained. Then, text detection is performed on the note text to obtain a first set of brand words; that is, brand words are detected first in the note text. Next, brand word detection is performed on the note image to obtain at least one second set of brand words; that is, brand word detection is performed on the note image again. This is achieved through multimodal detection. Furthermore, subject recognition is performed on the note image to obtain h categories for h subjects. Finally, based on the h categories of the h subjects, the first set of brand words and at least one second set of brand words are fused to obtain z target brand words. Therefore, when a user's screenshot behavior is detected, this application uses a multimodal detection method to detect brand words, thereby comprehensively detecting various types of brand words and improving the coverage of brand word detection. In addition, in order to avoid the situation where a single modality is not friendly to the detection of a certain brand word, this application combines the category of the subject and merges the brand words detected by multimodality to output the final brand words, so that the final target brand words are consistent in the dimensions of multimodality and subject category, thereby improving the detection accuracy of brand words. Attached Figure Description

[0080] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0081] Figure 1 A flowchart illustrating a method for obtaining brand terms provided in an embodiment of this application;

[0082] Figure 2 This application provides an embodiment of a method for obtaining note information based on a user's screenshot behavior.

[0083] Figure 3 A schematic diagram of the structure of a model provided in an embodiment of this application;

[0084] Figure 4 A schematic diagram illustrating the acquisition of brand terms as provided in an embodiment of this application;

[0085] Figure 5 This is a schematic diagram illustrating a brand recommendation for a user, provided as an embodiment of this application.

[0086] Figure 6 A schematic diagram illustrating the construction of a user profile for an embodiment of this application;

[0087] Figure 7Schematic structural diagram of a brand term acquisition device provided by an embodiment of the present application;

[0088] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0089] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0090] The terms "comprising" and "having" and any variations thereof appearing in the specification, claims and drawings of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. In addition, the terms "first", "second" and "third", etc. are used to distinguish different objects, rather than to describe a specific order.

[0091] First of all, for the convenience of understanding the brand term acquisition method of the present application, a brand term library is first constructed, which contains reference brand terms of each brand. And, the brand term library contains multiple reference brand terms for each brand, where the multiple reference brand terms include the full Chinese name, full English name, Chinese abbreviation, English abbreviation, alias, etc. of the brand. For example, for the brand Skechers, all types of reference brand terms of the brand are pre-constructed in the brand term library, that is, it includes "SKECHERS" (full English name), "斯凯奇" (full Chinese name), "S" (abbreviation).

[0092] It should be further noted that for each brand, a structured relationship between each brand and a category, that is, a key-value (k-v) structure, is also pre-constructed in the brand vocabulary. Here, the category is a coarse-grained category of the brand. For example, for the brand Skechers, the coarse-grained category of this brand is set as fashion, and the pre-constructed structured relationship is {Skechers}-{fashion}. The reason for constructing the structured relationship between the brand and the category is that when performing brand word matching later, while matching the corresponding brand word, the coarse-grained category corresponding to the brand word can also be determined, so as to determine the target confidence level of the brand word. The specific process of determining the target confidence level of the brand word based on the coarse-grained category of the brand word will be introduced in detail later, and will not be described in detail here.

[0093] Furthermore, as described above, there are multiple brand words for each brand. For example, the full Chinese name, the full English name, the Chinese abbreviation, the English abbreviation, the alias, and so on. Therefore, for each brand, a brand word can be selected from the multiple brand words of each brand as the main brand word. For example, for "SKECHERS" (full English name), "斯凯奇" (full Chinese name), "S" (abbreviation), "斯凯奇" can be selected as the main brand word. This application does not limit the method of selecting the main brand word. The reason for setting the main brand word is to map the brand words detected by different detection methods (i.e., different modalities) to the dimension of the main brand word, so as to facilitate the fusion of the brand words detected by different modalities. The process of mapping the brand word to the main brand word and the process of fusing the brand words will be introduced in detail later, and will not be described in detail here.

[0094] See Figure 1 , Figure 1 which is a schematic flowchart of a method for obtaining brand words provided by an embodiment of this application. This method is applied to a brand word acquisition device. The method includes but is not limited to the following steps:

[0095] 101: Obtain note information, where the note information includes a note picture and note text.

[0096] Optionally, when the brand word acquisition device detects a preset operation of the user, the note information can be obtained.

[0097] For example, the preset operation can be a screenshot action. That is, when a user's screenshot action is detected, the system responds to the user's screenshot action and retrieves the note information corresponding to that screenshot action. For example, the preset operation can also be triggered when the number of times a note image is viewed exceeds a threshold. For instance, when the brand keyword acquisition device detects that a user's number of views of a certain note image exceeds a threshold, it can retrieve the note information corresponding to that note image. For example, the preset operation can also be triggered when the viewing duration of a note image exceeds a threshold, and so on. This application does not limit the type of preset operation, as long as it can trigger the brand keyword acquisition device to retrieve note information. This application mainly uses the user's screenshot action as an example for explanation; other preset operations are similar and will not be described further.

[0098] For example, a user can browse notes under various note categories on the brand keyword acquisition device's application. Notes within each category can be posted by other users or by the user themselves. Note categories can include beauty, fashion, home decor, relationships, etc., and each category contains multiple notes, each containing a note image and / or text. When a user is interested in a note within a specific category, they can tap the note to view its images. While viewing a note image, the user can take a screenshot. When a user takes a screenshot, the brand keyword acquisition device detects this action and retrieves the corresponding note information. Optionally, the brand keyword acquisition device, in response to the screenshot action, determines the image ID of the note image and retrieves the corresponding note information based on that image ID. Specifically, when a user takes a screenshot of a note image, the brand keyword acquisition device can determine, based on the user's browsing history, which note within which note category the user wants to screenshot, that is, determine the image ID of the note image the user wants to screenshot, and then retrieve the note information corresponding to the screenshot behavior from the database based on the image ID.

[0099] For example, such as Figure 2As shown, when browsing notes, if a user is interested in note 2 under the "Fashion" category, they can click on note 2 to enter and view the note images within it. When the user views the third note image in note 2 and takes a screenshot (i.e., a screenshot action occurs), the brand keyword acquisition device responds to the user's screenshot action, determines that the user took a screenshot of the third note image in note 2 under the "Fashion" category, and then retrieves the note information corresponding to this screenshot action from the database, namely, the third note image in note 2 under the "Fashion" category, and the note text corresponding to that third note image.

[0100] 102: Perform subject recognition on the note image to obtain h subjects and h categories, with a one-to-one correspondence between the h subjects and h categories.

[0101] It should be noted that brand keywords are generally related to products. However, some notes in certain categories are purely for knowledge dissemination or other promotional purposes and do not involve products. For example, many notes in the "travel" category are used to promote scenery and attractions. These notes do not involve products and therefore do not involve brand keywords. In other words, these notes are of no value for brand keyword identification. Therefore, before performing subject recognition on note images, the note information can be valued by filtering its value, i.e., determining whether the note information is valuable, or whether it involves products. Specifically, note images and note text can be identified to determine whether the user's screenshot behavior is product-oriented and whether the note image and / or note text involve products. If at least one of the note image and note text involves a product, the note information is considered valuable, and the brand keyword acquisition device will then proceed with the subsequent brand keyword acquisition process. If neither the note image nor the note text involves a product, the note information is considered of no value and is skipped without brand keyword detection. This application uses a note information containing products as an example for illustration.

[0102] For example, object detection is performed on the note image to obtain multiple candidate boxes. Optionally, the note image is input into an object detection model, which performs object detection on the note image to obtain multiple candidate boxes and the position information of each candidate box, namely the center pixel coordinates of each candidate box in the note image, and the width and height of the candidate box in the note image, i.e., (x0, y0, w0, h0), where x0 and y0 are the pixel coordinates of the center pixels, w0 is the width, and h0 is the height. Then, based on the position information of each candidate box, subject detection is performed on the multiple candidate boxes to obtain h subject boxes. That is, the position information of each candidate box is input into the subject detection model to obtain the probability that each candidate box belongs to a subject box, that is, the probability that the object in the candidate box belongs to a subject. When the probability of a candidate box is greater than a threshold, the candidate box is regarded as a subject box, and h subject boxes are selected from the multiple candidate boxes. Here, h is an integer greater than or equal to 0. When h = 0, it means that there is no subject in the note image. This application does not consider this case much. This application mainly considers the case where subject boxes are detected.

[0103] It should be noted that the reason for filtering candidate bounding boxes is that object detection only detects objects in the note image and does not consider whether each object is the main subject. For brand keyword extraction, only brand keywords related to the main subject need to be extracted. Therefore, filtering by main subject bounding boxes avoids focusing on useless targets, improving the efficiency of brand keyword acquisition.

[0104] For example, if a note image is used to promote toothpaste, but a user accidentally includes a hair clip on a table in the photo, and another user is interested in the image and takes a screenshot, the hair clip will be listed as a target after target detection. However, the user is not interested in the hair clip; in other words, the user will not pay attention to the hair clip brand. Therefore, this application filters out the hair clip target before obtaining brand keywords, avoiding the processing of invalid information and improving the efficiency of brand keyword acquisition.

[0105] Furthermore, after obtaining h subject frames, the subjects in each subject frame are classified to obtain the category corresponding to the subject in each subject frame, thus obtaining h categories corresponding to h subjects.

[0106] Specifically, the content selected by each subject frame in the note image is input into the classification model to obtain the category corresponding to the subject in each subject frame. It should be noted that this application can perform single-level or multi-level classification when classifying the subject in each subject frame, i.e., outputting a single-level category or a multi-level category. This application mainly uses the output of a multi-level category as an example for illustration. Therefore, the category corresponding to each subject in this application is a multi-level category. The lower-level category is a subcategory of the higher-level category.

[0107] Specifically, when outputting multi-level categories, the classification model can be used to extract features from the content selected by each subject box in the note image to obtain image features. Based on these image features, the probability that the subject of each subject box falls into each preset category under each level category is obtained. The preset category with the highest probability is taken as the category under that level category, thereby obtaining the multi-level category corresponding to the subject of each subject box.

[0108] For example, when the subject selected in the subject box is shampoo, the classification model can output the category of the subject as beauty_hair care and cleaning agent_hair cleaning agent_shampoo, where beauty is the first-level category, hair care and cleaning agent is the second-level category, hair cleaning agent is the third-level category, and shampoo is the fourth-level category.

[0109] As can be seen, in this embodiment, during target detection, targets are not directly classified; that is, the category of each target is not output. Instead, the category of each subject is classified only after the subject is selected. This avoids classifying non-subject targets, reducing the classification burden. Furthermore, classifying subjects after determining the subject bounding box allows for fine-grained classification, outputting multi-level categories for each subject instead of a single coarse category. For example, if the target is shampoo, classifying it during target detection would only output the category "hair care products," failing to achieve fine-grained classification. In this application, classification after subject detection outputs a four-level category, achieving fine-grained classification. After achieving fine-grained classification, subsequent brand word fusion can be based on these fine-grained categories, improving the detection accuracy of brand words.

[0110] In one possible implementation, after determining the category of each subject, the subjects can be sorted to obtain the order of each subject among the h subjects. Specifically, the category of each subject and the position information of the subject bounding box of each subject are input into the subject sorting model to obtain the confidence score of each subject; based on the confidence score of each subject, the h subjects are sorted to obtain the order of each subject, where the order of each subject is used to characterize the credibility of each subject belonging to a real subject. It should be noted that after sorting the subjects, the subsequent brand word fusion can be performed according to the order of the subjects. The specific fusion process will be described later and will not be described in detail here.

[0111] It should be noted that the above-mentioned object detection model, subject detection model, classification model, and subject ranking model are all pre-trained models. The training process can be supervised, unsupervised, or semi-supervised. No restrictions are placed on the training method.

[0112] 103: Perform text detection on the notes to obtain the first set of brand keywords.

[0113] For example, entity recognition is performed on the note text to obtain 'a' first candidate brand words, which means performing NER detection on the note text to obtain 'a' first candidate brand words. Here, 'a' is an integer greater than or equal to 0. It should be understood that when 'a' = 0, no brand words are detected by NER detection, and therefore no first brand word set is generated; or it can be understood that the first brand word set is an empty set.

[0114] Specifically, the note text is subjected to entity word recognition to obtain at least one entity word. Then, each entity word is matched with each reference brand word in the brand word library to obtain the matching degree with each reference brand word; reference brand words with a matching degree greater than a threshold are selected as candidate brand words; after matching all at least one entity word, a first candidate brand words can be obtained.

[0115] Furthermore, a primary brand term mapping is performed on each first-candidate brand term to obtain the corresponding brand term. Finally, the brand terms corresponding to each first-candidate brand term are combined into the aforementioned first-brand term set. Since the brand term corresponding to each first-candidate brand term can be understood as the primary brand term corresponding to each first-candidate brand term, the brand terms in the first-brand term set can also be called primary brand terms. They are essentially the same and do not need to be distinguished.

[0116] Understandably, by matching each entity word with each reference brand word in the brand terminology library and obtaining a reference brand word with a matching degree greater than the threshold, the coarse-grained category corresponding to the reference brand word can also be obtained. This allows the coarse-grained category of each brand word in the first brand terminology set to be determined, and a structured relationship between each brand word in the first brand terminology set and the coarse-grained category to be established.

[0117] 104: Perform brand keyword detection on the note images to obtain at least one set of secondary brand keywords.

[0118] In one embodiment of this application, a logo detection is performed on the note image to obtain the at least one set of second brand words.

[0119] For example, based on logo detection, brand word detection is performed on the note image to obtain at least one sixth candidate brand word. That is, the logo in the note image is identified, and based on the identified logo, at least one sixth candidate brand word is determined. Then, each sixth candidate brand word is fuzzily matched with reference brand words in the brand word library to obtain b second candidate brand words. Here, b is an integer greater than or equal to 0. It should be understood that when b = 0, no brand word is detected by logo detection, and in this case, no second brand word set is obtained through logo detection; or it can be understood that the second brand word set obtained through logo detection is an empty set.

[0120] Then, a primary brand term is mapped to each second candidate brand term to obtain the corresponding brand term (i.e., the primary brand term). Finally, at least one brand term corresponding to each of the b second candidate brand terms is combined into at least one set of second brand terms. Similarly, the brand terms in the set of second brand terms can also be called primary brand terms; they are essentially the same and do not need to be distinguished.

[0121] It is understood that in this embodiment, brand word detection is performed on the note image only through logo detection. Therefore, the at least one set of second brand words obtained in this embodiment is a set of second brand words, and this set of second brand words is the set of second brand words obtained by combining at least one brand word corresponding to b second candidate brand words.

[0122] Similarly, after performing fuzzy matching with the brand term library, we can obtain the coarse-grained category of each second candidate brand term. This allows us to determine the coarse-grained category of each main brand term in the second brand term set and establish a structured relationship between each main brand term in the second brand term set and its coarse-grained category.

[0123] In another embodiment of this application, optical character recognition (OCR) is performed on the note image to obtain the at least one set of second brand words.

[0124] For example, optical character recognition (OCR) is performed on the note image to obtain c third candidate brand words. Here, c is an integer greater than or equal to 0. It should be understood that when b = 0, no brand words are detected by OCR, and in this case, OCR will not yield a second set of brand words; or it can be understood that the second set of brand words obtained by OCR is an empty set.

[0125] For example, text extraction is performed on the note image using OCR technology to obtain the text information in the note image. Optionally, the text information is matched with a brand terminology database to obtain d fourth candidate brand terms. Here, d is an integer greater than or equal to 0. It should be understood that when d = 0, it indicates that the text information failed to match the brand terminology database, and no fourth candidate brand term was obtained.

[0126] Specifically, keywords are extracted from the text information to obtain multiple keywords. Then, each keyword is matched with each reference brand word in the brand terminology library to obtain the fourth candidate brand word corresponding to each keyword, that is, d fourth candidate brand words are obtained.

[0127] Similarly, after matching the text information with the brand terminology database to obtain each fourth candidate brand term, we can also obtain the coarse-grained category corresponding to each fourth candidate brand term.

[0128] Optionally, the text information is fuzzily matched with the above h categories to obtain e fifth candidate brand words. Here, e is an integer greater than or equal to 0. It should be understood that when e = 0, it indicates that the fuzzy match between the text information and the above h categories failed, and no fifth candidate keyword was obtained.

[0129] Specifically, the text information is fuzzily matched with each of the h categories to obtain successfully matched keywords. More specifically, the text information is fuzzily matched with each level of category of each subject to identify keywords in the text information that match all h categories. These keywords are considered successfully matched keywords. Category consistency means that if a keyword matches a certain level of category of a subject, it is considered a keyword that matches all h categories. Then, a predetermined number of characters preceding the keywords in the text information are used as a fifth candidate brand word, resulting in e fifth candidate brand words.

[0130] For example, if the text message is "My Vidal Sassoon shampoo is not yet finished", and the h categories contain shampoo, then when performing fuzzy matching, the keyword that can be successfully matched is "shampoo". The first two characters of this keyword can be used as the brand name, thus obtaining a fifth candidate brand name: "Vidal Sassoon".

[0131] It should be noted that since the fifth candidate brand name is obtained through fuzzy matching of categories, the coarse-grained category of each fifth candidate brand name is the category that matches the h categories. For example, the brand name "Vidal Sassoon" matched above has the coarse-grained category of "shampoo".

[0132] It should be noted that the reason for performing fuzzy matching between the text information and the aforementioned h categories is that OCR technology simply extracts text information and is insensitive to brand words. Therefore, the segmented keywords are also insensitive to brand words; that is, it does not intentionally segment brand words as keywords. This results in the incomplete segmentation of all brand words. For example, if the text information is "My Vidal Sassoon shampoo is not finished," during keyword segmentation, "Vidal Sassoon" will not be considered as a keyword. Thus, when matching with the brand word database, not all brand words in the text information will be extracted. Then, by performing fuzzy matching between the text information and the h categories, when any of the h categories contains "shampoo," the keyword "shampoo" in the text information can be matched, along with the two characters before "shampoo," namely "Vidal Sassoon." It can be seen that by performing fuzzy matching between the text information and the aforementioned h categories, brand word recall can be expanded, allowing for the extraction of richer brand words from the text information. This results in richer and more comprehensive candidate brand words, improving the comprehensiveness of brand word extraction by OCR.

[0133] Finally, the d fourth candidate brand words and e fifth candidate brand words are all used as the c third candidate brand words. Further, a main brand word mapping is performed on each third candidate brand word to obtain the corresponding brand word (i.e., the main brand word). Finally, the brand words corresponding to each third candidate brand word are combined into at least one set of second brand words as described above.

[0134] It should be noted that since each third candidate brand term has a corresponding coarse-grained category, a structured relationship can be constructed between each third candidate brand term and its coarse-grained category.

[0135] It is understood that in this embodiment, only OCR recognition is performed on the note image, so the number of the at least one second brand word set is one, and the second brand word set is the second brand word set obtained by combining the brand words corresponding to each third candidate brand word.

[0136] In another embodiment of this application, logo detection and optical character recognition can be performed simultaneously on the note image to obtain at least one second brand word set.

[0137] Specifically, performing logo detection on the note image yields a second set of brand words. The process of obtaining this second set of brand words is similar to the process of using logo detection alone for brand word detection described above, and will not be described again. Performing OCR detection on the note image yields another second set of brand words. The process of obtaining this second set of brand words is similar to the process of using OCR detection alone for brand word detection described above, and will not be described again.

[0138] It is understood that in this embodiment, two detection methods are used to detect brand words in the note image simultaneously, so the number of the above-mentioned at least one second brand word set is two.

[0139] It should be noted that in practical applications, whether to use logo detection or OCR detection alone to obtain at least one set of secondary brand words, or to use both logo detection and OCR detection simultaneously, can be preset or selected based on the note category. For example, when the note category is "fashion," one can choose to use both logo detection and OCR detection simultaneously to obtain at least one set of secondary brand words; when the note category is not "fashion," one can choose to use only logo detection to obtain at least one set of secondary brand words. This application mainly uses the simultaneous use of logo detection and OCR detection to obtain brand words from note images as an example, but it is not limited to using only logo detection or OCR detection.

[0140] 105: Based on h categories of h subjects, merge the first brand term set and at least one second brand term set to obtain z target brand terms. Where z is an integer greater than or equal to 0.

[0141] For example, firstly, determine m first target brand words from a first set of brand words and at least one second set of brand words. That is, determine brand words that exist in at least two sets of brand words, one first set and one second set. Use these brand words as the first target brand words, resulting in m first target brand words. Here, m is an integer greater than or equal to 0. It should be understood that when m = 0, it means that the brand words detected under different modalities are not the same. Since the first target brand words exist in at least two sets of brand words, meaning there are at least two detection methods, if the note information contains the first target brand word, it can be confidently concluded that the user's screenshot behavior necessarily includes the first target brand word. Therefore, the m first target brand words can be directly output.

[0142] The above describes the simultaneous recall of m primary target brand keywords using three detection methods, i.e., three modalities. However, some brand keywords that only exist in each modality may also be brand keywords included in the note information. Therefore, the following describes a method for recalling brand keywords using a single modality.

[0143] For example, m first target brand words are removed from the first brand word set and at least one second brand word set to obtain a third brand word set corresponding to the first brand word set and a fourth brand word set corresponding to each second brand word set.

[0144] It should be noted that since the first target brand term exists in at least two brand term sets, not every brand term set necessarily contains the first target brand term. Therefore, when removing the m first target brand terms, if a certain brand term set contains the m first target brand terms, then the m first target brand terms are removed from that brand term set to obtain a new brand term set; if not, the brand term set can be directly retained as the new brand term set.

[0145] For example, for a third set of brand words, where the coarse-grained category of each brand word in the third set of brand words is known, the brand words in the third set of brand words can be filtered according to h categories to obtain n second target brand words.

[0146] Specifically, at least one fourth-category brand term is obtained from the third brand term set, where the h categories contain the category corresponding to each fourth target brand term. That is, brand terms that match the h categories are selected from the third brand term set as fourth target brand terms. Specifically, the coarse-grained category of each brand term in the third brand term set is obtained; then, the coarse-grained category of each brand term is compared with each category. If the coarse-grained category of a brand term matches a certain category, then that brand term is selected as the fourth target brand term. Specifically, comparing the coarse-grained category of each brand term with each category means comparing the coarse-grained category of each brand term with each level of each category. If it matches a certain level of category, then the coarse-grained category of the brand term is determined to match that category.

[0147] It should be noted that when comparing with each of the h categories, the comparison can be performed separately with each category according to the order of the entities corresponding to each category as determined above. This is because entities ranked higher are more likely to be genuinely relevant. If the entities ranked higher can be successfully compared, there is no need to compare with the categories ranked lower, thus improving the efficiency of brand keyword acquisition.

[0148] For example, if a subject's category is Other_Hydration & Cleaning Agents_Oral Care_Toothpaste, and a brand name is "Haolai," and the coarse-grained category corresponding to "Haolai" is "Oral Care," it can be seen that the coarse-grained category of the brand name is consistent with the third-level category of the subject. Therefore, it is determined that the brand name is consistent with the subject's category, and this brand name can be used as the fourth target brand name.

[0149] Furthermore, remove x fourth target brand terms from the third brand term set to obtain the fifth brand term set. That is, remove the x fourth target brand terms from the third brand term set to obtain the fifth brand term set. Here, x is an integer greater than or equal to 0. When x = 0, it means that the category of any brand term in the third brand term set is inconsistent with the aforementioned h categories.

[0150] Furthermore, a weighted processing operation is performed based on the category corresponding to each brand word in the fifth brand word set to obtain the target confidence score for each brand word in the fifth brand word set. Based on the target confidence scores of each brand word in the fifth brand word set, y fifth target brand words are obtained. Specifically, brand words with a target confidence score greater than a first threshold are selected as fifth target brand words. Here, y is an integer greater than or equal to 0.

[0151] For example, for each set of fourth brand words, brand words are filtered according to h categories to obtain k third target brand words corresponding to each set of fourth brand words. Here, k is an integer greater than or equal to 0. It should be noted that for the first brand word sets detected by different detection methods, the number of k third target brand words obtained by filtering the fourth brand word set can be the same or different. That is, for the first brand word sets detected by different detection methods, the value of k can be the same or different.

[0152] Specifically, obtain q sixth target brand words from each fourth brand word set, where the h categories contain the category corresponding to each sixth target brand word. That is, select brand words from each fourth brand word set that match the h categories as fifth target brand words to obtain q sixth target brand words. The selection process is similar to the process of selecting fourth target brand words from the third brand word set described above, and will not be described again. Here, q is an integer greater than or equal to 0.

[0153] Further, remove q sixth target brand words from each fourth brand word set, that is, remove q sixth target brand words from the fourth brand word set to obtain the sixth brand word set corresponding to the fourth brand word set. Then, perform a weighted processing operation according to the category corresponding to each brand word in the sixth brand word set to obtain the target confidence score of each brand word in the sixth brand word set. Finally, based on the target confidence score of each brand word in the sixth brand word set, determine r seventh target brand words corresponding to each fourth brand word set, that is, brand words in the sixth brand word set with a target confidence score greater than the second threshold are taken as seventh target brand words, thus obtaining the r seventh target brand words.

[0154] Finally, the q sixth target brand words in each fourth brand word set and the r seventh target brand words corresponding to each fourth brand word set are merged (i.e., combined) to obtain the k third target brand words corresponding to each fourth brand word set.

[0155] The following section uses brand keyword A as an example to detail how to perform the weighted processing operation. Brand keyword A can be any one of the fifth brand keyword set or any one of the sixth brand keyword sets. For brand keyword A, firstly, the penalty coefficient corresponding to brand keyword A is determined based on its category. Then, based on the brand keyword set in which brand keyword A belongs, the preset penalty item corresponding to brand keyword A is obtained. Finally, based on the penalty coefficient, the preset penalty item, and the confidence level of brand keyword A, the target confidence level corresponding to brand keyword A is determined.

[0156] First, it should be clarified that for NER detection, the confidence score of the detected brand words is the probability that the detected entity word belongs to the brand word when matched with the brand word, which is a value of 0 or 1. Similarly, for OCR detection, the confidence score of the detected brand words is also the probability that the detected entity word belongs to the brand word when matched with the brand word, which is also a value of 0 or 1. The confidence score of the brand words detected by Logo is the confidence score of the brand words detected by the Logo detection model.

[0157] Specifically, for each detection mode, brand keyword detection is relatively more favorable for certain categories of brand keywords, meaning that the recognition accuracy is higher when identifying these categories. For example, logo detection is more favorable for clothing-related brand keywords, while NER detection is more favorable for beauty-related brand keywords. Therefore, different penalty coefficients can be set for different categories of brand keywords detected for each detection mode.

[0158] Optionally, the penalty coefficient can be 1 or -1. Of course, in practical applications, other values ​​can also be set, as long as the following requirements are met: through the penalty coefficient, when the category of a brand word is a brand word that is relatively friendly to this detection mode, the confidence of the brand word is increased to obtain the target confidence of the brand word; when the category of a brand word is a brand word that is not friendly to this detection mode, the confidence of the brand word is decreased to obtain the target confidence of the brand word.

[0159] Optionally, for logo detection, the penalty coefficient for brand words in the first category is set to 1, and the penalty coefficient for brand words not in the first category is set to -1, where the first category can be clothing; Optionally, for NER detection, the penalty coefficient for brand words in the second category is set to 1, and the penalty coefficient for brand words not in the second category is set to -1, where the second category can be beauty; Optionally, for OCR detection, the penalty coefficient for brand words in the third category is set to 1, and the penalty coefficient for brand words not in the third category is set to -1, where the third category is beauty.

[0160] Therefore, the penalty coefficient for brand term A can be determined based on the category of brand term A (i.e., the coarse-grained category of brand term A).

[0161] Furthermore, different penalty items can be set for different detection modes, and this application does not impose any restrictions on this. This application mainly uses setting different penalty items as an example for illustration. For example, the penalty item for OCR detection and NER detection is set to 1, and the penalty item for Logo detection is set to a value between 0 and 1.

[0162] For example, the target confidence level of brand term A can be expressed by formula (1):

[0163] P t =P i +μ*σ Formula (1);

[0164] Among them, P t For the target confidence level of brand keywords, P i σ represents the confidence level of brand term A, μ represents the penalty coefficient corresponding to brand term A, and σ represents the preset penalty item corresponding to brand term A.

[0165] It should be noted that when brand word A belongs to the fifth set of brand words or the sixth set of brand words obtained through OCR detection, P i The value is either 0 or 1; when brand word A belongs to the brand word in the fifth brand word set or the brand word in the sixth brand word set obtained through Logo detection, P... i The value is based on the confidence level that the brand word A belongs to the brand word detected by the logo, and the value is between 0 and 1.

[0166] Finally, after obtaining m first target brand terms through multimodal synthesis, n second target brand terms through unimodal output, and k third target brand terms corresponding to each fourth brand set, the m first target brand terms, n second target brand terms, and k third target brand terms corresponding to each fourth brand term set are merged to obtain z target brand terms. That is, by merging and deduplicating the m first target brand terms, n second target brand terms, and k third target brand terms corresponding to each fourth brand term set, the aforementioned z target brand terms can be obtained.

[0167] In one embodiment of this application, it should be noted that some note information only contains note images and does not contain note text; or, some note information contains incomplete note text, making accurate brand word detection impossible; or, when performing brand word detection, the note text is not used for brand word detection; and so on. In general, after detecting a user's screenshot behavior, brand words can be obtained solely through logo detection and OCR detection, without needing to focus on the note text.

[0168] For example, in response to a user's screenshot action, a note image corresponding to the screenshot action is obtained; subject recognition is performed on the note image to obtain h subjects and h categories, with each of the h subjects and h categories corresponding one-to-one; logo detection is performed on the note image to obtain a second set of brand words; OCR detection is performed on the note image to obtain another second set of brand words; based on the h subjects and h categories, the two second set of brand words are fused to obtain z target brand words.

[0169] It should be noted that the process of obtaining the second set of brand words can be referred to the process of obtaining the second set of brand words described above, and will not be repeated here; and the process of merging the two sets of second brand words can be referred to the process of merging brand word sets described above, and will not be repeated here either.

[0170] As can be seen, in this embodiment, when a user's screenshot behavior is detected, the note image corresponding to the screenshot behavior can be obtained. The brand word recognition and fusion can be achieved in a multimodal manner based solely on the note image. Therefore, while ensuring the accuracy of brand word recognition, the recognition efficiency of brand words can also be improved.

[0171] In one embodiment of this application, the note information also includes note categories.

[0172] Optionally, the note category can be used during text detection of the note text, i.e., during NER detection of the note text, to provide prior information for the text detection, so that the brand words in the first set of brand words obtained by the text detection can be matched with the note category. Specifically, using the note category can provide prior information (guidance information) for the identification of brand words in advance, that is, to indicate that the identified brand words should fall under this note category, thereby avoiding the identification of some brand words that do not belong to this note category, thus improving the accuracy of brand word identification during text detection.

[0173] Optionally, the note category can also provide prior information when performing brand word detection on note images, so that the brand words in each obtained second brand word set match the note category. Optionally, when performing logo detection on note images, this note category can be used. This way, when performing brand word recognition based on the logo, the identified brand words can be placed under this note category, thus avoiding the identification of brand words that do not belong to this note category, thereby improving the brand word recognition accuracy during logo detection. Optionally, when performing OCR detection on note images, this note category can be used. This way, when the text information obtained through OCR detection is fuzzily matched with the brand word database, the matched brand words can be placed under this note category, thus avoiding the identification of brand words that do not belong to this note category, thereby improving the brand word recognition accuracy during OCR detection.

[0174] Optionally, the note category can also provide prior information when performing subject recognition on the note image, so that the obtained subject category matches the note category. Specifically, when performing object detection on the note image, the note category can be used to guide the identified object to match the note category, thereby improving the accuracy of object detection. Furthermore, when performing topic recognition on the object, the note category can also be used, so that the identified subject matches the note category, thereby improving the accuracy of subject recognition. Furthermore, when classifying the subject, the note category can also be used, so that the classified category matches the note category, thereby improving the accuracy of subject classification.

[0175] In one embodiment of this application, the process of obtaining target brand keywords can be implemented through a model. The process of obtaining target brand keywords described above is explained below with reference to the specific structure of the model.

[0176] See Figure 3 , Figure 3 This is a schematic diagram of the structure of a model provided in an embodiment of this application. Figure 3As shown, the model includes an optional note value filtering module, a brand keyword detection module, a subject recognition module, a brand keyword expansion and recall & structuring module, and a brand keyword fusion module. The brand keyword detection module includes a NER detection model, a Logo detection model, and an OCR detection model. The subject recognition module includes an object detection model, a subject detection model, a classification model, and a subject ranking model (optionally). The brand keyword expansion and recall & structuring module includes NER expansion and recall & structuring modules, Logo expansion and recall & structuring modules, and OCR expansion and recall & structuring modules.

[0177] For example, when a user's screenshot action is detected, the system responds by acquiring the corresponding note information, namely the note image, note text, and note category (optional). As mentioned above, the note category may or may not participate in the subsequent brand word recognition process. The following explanation uses the example of the note category not participating in the brand word recognition process. Then, the note information is input into the note value filtering module for value filtering to determine whether the note information has transaction value, i.e., whether the note information contains product information. If it does, the brand word acquisition method of this application is executed to output the brand word; otherwise, the brand word is not output. It should be noted that in practical applications, the note value filtering module may not be designed. That is, after acquiring the note information corresponding to the user's screenshot action, no value filtering is performed on the note information, which is equivalent to assuming that every note information has value.

[0178] Furthermore, when the note information is determined to be valuable, the note image and note text are input into the brand word detection module. The NER detection model performs brand word entity recognition on the note text to obtain 'a' first candidate brand words; the Logo detection model performs brand word detection on the note image to obtain 'b' second candidate brand words; and the OCR detection model performs brand word detection on the note image to obtain 'c' third candidate brand words. The process of each model detecting brand words can be found in steps 103 and 104 above, and will not be described further.

[0179] Further, the first candidate brand words (a) are input into the NER (Network Response) expansion and structuring module. For each first candidate brand word, expansion and recall are performed, i.e., main brand word mapping, to obtain the main brand word corresponding to each first candidate brand word. A structured relationship is then established between the main brand word and its coarse-grained category for each mapped main brand word, resulting in a set of first brand words. Similarly, the second candidate brand words (b) are input into the Logo expansion and structuring module. For each second candidate brand word, expansion and recall are performed, i.e., main brand word mapping, to obtain the main brand word corresponding to each second candidate brand word. The main brand term is assigned to each candidate brand term, and a structured relationship is established between the main brand term and its coarse-grained category for each candidate brand term obtained through mapping, resulting in a set of second brand terms. Then, c third candidate brand terms are input into the OCR recall & structuring module. For each third candidate brand term, recall is performed to obtain the corresponding main brand term. A structured relationship is then established between the main brand term and its coarse-grained category for each candidate brand term obtained through mapping, resulting in another set of second main brand terms.

[0180] In addition, the note image is input into the subject recognition module for subject identification. Specifically, the note image is first input into the object detection model to obtain multiple candidate boxes and the position information of each candidate box; then, the position information of each candidate box is input into the subject detection model to filter out h subject boxes from the multiple candidate boxes; then, the content selected by each subject box in the note image is input into the classification model to obtain the category of the subject selected by each subject box, that is, h categories of h subjects. Then, the position information of each subject box and the category of the subject in each subject box can be input into the subject ranking model to obtain the order of each subject box, so that when performing brand word fusion in the subsequent brand word fusion module, the fusion can be performed based on the order of each subject box. Of course, in practical applications, the subject ranking model can also be omitted.

[0181] Furthermore, the main brand words from the first brand word set, the main brand words from the two second brand word sets, and h categories from at least one theme are all input into the brand word fusion module for fusion to obtain z target brand words. The fusion process of brand words can be referred to in step 104 above, and will not be described again.

[0182] The following example illustrates the brand term acquisition method used in this application.

[0183] See Figure 4 , Figure 4 This is a schematic diagram illustrating how to obtain brand words according to an embodiment of this application.

[0184] like Figure 4As shown in the figure, the note picture of the user screenshot includes three main subjects. First, in response to the user's screenshot behavior, the picture ID of the note picture to be screenshot by the user is determined first, that is, beauty_personal care_62b454bf000000000102e32c; then, based on this picture ID, the note picture and the corresponding note text are retrieved from the database. This note picture corresponds to the note text. Then, the brand words of the note picture are detected through the Logo, and the brand word obtained is CHANEL. It can be seen that the Logo detection is not very friendly to the brand words of the beauty category, so the brand words are not accurately detected; since OCR is more friendly to the brand words of the beauty category, the brand words of the note picture are detected through OCR, and the brand word DRALIE is obtained, that is, the brand words in the note picture can be accurately recognized; since there is no note text, no brand words will be output when using NER for brand word detection; at the same time, the main subjects of the note picture are recognized, and it can be recognized that there are three main subjects in the note picture, and the categories of these three main subjects are all other_laundry detergents_dental care_toothpaste.

[0185] Then, the main brand word mapping is performed on the brand word CHANEL detected by the Logo, and the main brand word "Chanel" can be obtained, and the established structured relationship is: "Chanel" - "dressing"; the main brand word mapping is performed on the brand word detected by OCR, and the main brand word "Hora" can be obtained, and the established structured relationship is: "Hora" - "laundry detergents".

[0186] Furthermore, first perform multi-modal brand word fusion, and there is no corresponding target brand word; then, perform single-modal brand word input. It can be seen that since the coarse-grained category of the brand word "Hora" detected by OCR is the same as the category of the main subject, the brand word "Hora" can be output when comparing the categories; since the coarse-grained category of the brand word "Chanel" detected by the Logo is different from the category of the main subject, the "Chanel" will not be output when comparing the categories; finally, perform a weighting operation on the brand word "Chanel", and the target confidence level of the brand word "Chanel" is less than the threshold, so it is confirmed that the brand word "Chanel" is a misdetected brand word. Finally, as Figure 4 shown, the output target brand word is "Hora".

[0187] Combined with the brand word acquisition method provided in the embodiment of the present application, the application scenario of the present application is described.

[0188] In an embodiment of the present application, according to the z target brand words obtained above, brand recommendations are made for users. Please refer to Figure 5 , Figure 5 which is a schematic diagram of making brand recommendations for users provided in the embodiment of the present application.

[0189] like Figure 5 As shown, a user browses notes in an application on their device and takes a screenshot of a note image that interests them. The user device responds to the user's screenshot action by retrieving the note information corresponding to that screenshot. Then, the brand keyword retrieval method described above is executed on the note information to obtain z target brand keywords that the user is interested in. For example... Figure 5 As shown, the system retrieves the brands corresponding to the z target brand keywords and displays the visual interface of the user's device for that brand. When displaying the brand, it may show a brand link or image, etc.; this application does not limit the content displayed.

[0190] In another embodiment of this application, a user profile is constructed based on the z target brand keywords obtained above, that is, the z target brand keywords are used as a tag for the user. See also Figure 6 , Figure 6 This is a schematic diagram illustrating the construction of a user profile for a user, as provided in an embodiment of this application.

[0191] like Figure 6 As shown, a user browses notes in an application on their device and takes a screenshot of a note image that interests them. The user device responds to the user's screenshot action by retrieving the note information corresponding to that screenshot. Then, the brand keyword retrieval method described above is executed on the note information to obtain z target brand keywords that the user is interested in. Finally, as... Figure 6 As shown, each target brand keyword is used as a tag for users to build user profiles.

[0192] It should be noted that the brand keyword acquisition method of this application can also be applied to other scenarios, such as: providing users with corresponding shopping solutions based on target brand keywords, providing users with related notes based on target brand keywords, and recommending relevant stores to users based on target brand keywords, etc., without limiting the application scenarios of the brand keyword acquisition method of this application.

[0193] Please see Figure 7 , Figure 7 This is a schematic diagram of a brand name acquisition device provided in an embodiment of this application, as shown below. Figure 7 As shown, the brand keyword acquisition device includes an acquisition unit 701 and a processing unit 702;

[0194] in:

[0195] The acquisition unit 701 is used to acquire note information, wherein the note information includes note images and note text;

[0196] Processing unit 702 is used to perform subject recognition on the note image to obtain h categories of h subjects, wherein the h subjects and the h categories correspond one-to-one;

[0197] Text detection is performed on the notes to obtain a first set of brand keywords;

[0198] Brand word detection is performed on the notes image to obtain at least one second brand word set;

[0199] Based on the h categories of the h subjects, the first brand word set and the at least one second brand word set are merged to obtain z target brand words.

[0200] In one possible implementation, in fusing the first brand term set and the at least one second brand term set according to the h categories of the h subjects to obtain z target brand terms, the processing unit 702 is specifically used for:

[0201] Determine m first target brand words from the first brand word set and the at least one second brand word set, wherein each first target brand word exists in at least two brand word sets, the first brand word set and the at least one second brand word set;

[0202] Remove the m first target brand words from the first brand word set and the at least one second brand word set to obtain a third brand word set corresponding to the first brand word set, and a fourth brand word set corresponding to each second brand word set;

[0203] Based on the h categories, the third brand word set is filtered to obtain n second target brand words;

[0204] Based on the h categories, brand words are filtered for each fourth brand word set to obtain k third target brand words corresponding to each fourth brand word set;

[0205] The m first target brand words, the n second target brand words, and the k third target brand words corresponding to each fourth brand word set are merged to obtain the z target brand words.

[0206] In one possible implementation, in filtering the third brand term set according to the h categories to obtain n second target brand terms, the processing unit 702 is specifically used for:

[0207] Obtain x fourth target brand words from the third brand word set, wherein the h categories contain the category corresponding to each fourth target brand word;

[0208] Remove the x fourth target brand terms from the third brand term set to obtain the fifth brand term set;

[0209] A weighted processing operation is performed on each brand word in the fifth brand word set according to its corresponding category to obtain the target confidence level of each brand word in the fifth brand word set;

[0210] Based on the target confidence level of each brand word in the fifth brand word set, y fifth target brand words are obtained;

[0211] The x fourth target brand words and the y fifth target brand words are merged to obtain the n second target brand words.

[0212] In one possible implementation, in filtering brand words for each fourth brand word set according to the h categories to obtain k third target brand words corresponding to each fourth brand word set, the processing unit 702 is specifically used for:

[0213] Obtain q sixth target brand words from each fourth brand word set, wherein the h categories contain the category corresponding to each sixth target brand word;

[0214] Remove the q sixth target brand words from each fourth brand word set to obtain the sixth brand word set corresponding to each fourth brand word set;

[0215] A weighted processing operation is performed on each brand word in the sixth brand word set according to its corresponding category to obtain the target confidence level of each brand word in the sixth brand word set;

[0216] Based on the target confidence level of each brand word in the sixth brand word set, determine r seventh target brand words corresponding to each fourth brand word set;

[0217] The q sixth target brand words in each fourth brand word set and the r seventh target brand words corresponding to each fourth brand word set are merged to obtain the k third target brand words corresponding to each fourth brand word set.

[0218] In one possible implementation, in performing the weighted processing operation, the processing unit 702 is specifically used for:

[0219] Based on the category corresponding to brand word A, determine the penalty coefficient corresponding to brand word A, wherein brand word A is any one of the fifth brand word set or any one of each sixth brand word set;

[0220] Based on the brand word set in which brand word A is located, obtain the preset penalty item corresponding to brand word A;

[0221] The target confidence level of brand term A is determined based on the penalty coefficient corresponding to brand term A, the preset penalty item, and the confidence level of brand term A.

[0222] In one possible implementation, in performing text detection on the first note text to obtain a first brand term set, the processing unit 702 is specifically used for:

[0223] Entity recognition is performed on the note text to obtain a first candidate brand words;

[0224] Map each first-candidate brand term to a main brand term to obtain the brand term corresponding to each first-candidate brand term.

[0225] The brand words corresponding to each first candidate brand word are combined into the first brand word set.

[0226] In one possible implementation, in performing brand word detection on the note image to obtain at least one second set of brand words, the processing unit 702 is specifically used for:

[0227] Brand logo detection was performed on the note image to obtain b second candidate brand words;

[0228] Map each second candidate brand term to the main brand term to obtain the brand term corresponding to each second candidate brand term.

[0229] The brand words corresponding to each second candidate brand word are combined into at least one set of second brand words.

[0230] In one possible implementation, in performing brand word detection on the note image to obtain at least one second set of brand words, the processing unit 702 is specifically used for:

[0231] Optical character recognition was performed on the note image to obtain c third candidate brand words;

[0232] For each third candidate brand term, perform brand term mapping to obtain the brand term corresponding to each third candidate brand term;

[0233] The brand words corresponding to each third candidate brand word are combined into at least one set of second brand words.

[0234] In one possible implementation, in performing brand word detection on the note image to obtain at least one second set of brand words, the processing unit 702 is specifically used for:

[0235] Logo detection was performed on the note image to obtain b second candidate brand words;

[0236] Map each second candidate brand term to the main brand term to obtain the brand term corresponding to each second candidate brand term.

[0237] Optical character recognition was performed on the note image to obtain c third candidate brand words;

[0238] Map each second candidate brand term to a main brand term to obtain the brand term corresponding to each third candidate brand term.

[0239] Each second candidate brand word is combined into a second brand word set, and each third candidate brand word is combined into another second brand word set to obtain the at least one second brand word set.

[0240] In one possible implementation, in performing optical character recognition on the note image to obtain c third candidate brand words, the processing unit 702 is specifically used for:

[0241] Optical character recognition is performed on the note image to obtain text information;

[0242] The text information is matched with the brand terminology database to obtain d fourth candidate brand terms;

[0243] The text information is matched with the h categories to obtain e fifth candidate brand words;

[0244] The d fourth candidate brand words and the e fifth candidate brand words are used as the c third candidate brand words.

[0245] In one possible implementation, in performing subject recognition on the note image to obtain h subjects and h categories, the processing unit 702 is specifically used for:

[0246] Object detection is performed on the note image to obtain multiple candidate boxes;

[0247] Based on the position information of each candidate box, subject detection is performed on the multiple candidate boxes to obtain h subject boxes, wherein the target selected by each subject box is a subject;

[0248] The subject corresponding to each subject frame is classified to obtain h categories corresponding to each subject.

[0249] In one possible implementation, the note information further includes note categories; the note categories are used to provide prior information when performing text detection on the note text, so that brand words in the obtained first set of brand words match the note categories; the note categories are also used to provide prior information when performing brand word detection on the note images, so that brand words in each obtained second set of brand words match the note categories; the note categories are also used to provide prior information when performing subject recognition on the note images, so that the category of each obtained subject matches the note categories.

[0250] In one possible implementation, the processing unit 702 is further configured to:

[0251] Based on z target brand keywords, recommend brands to users.

[0252] In one possible implementation, the processing unit 702 is further configured to:

[0253] Based on z target brand keywords, build user profiles for users.

[0254] According to one embodiment of this application, Figure 7 The various units of the brand term acquisition device shown can be individually or entirely combined into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above units are based on logical function division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the cloud server may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0255] According to another embodiment of this application, the following can be achieved by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), a device capable of performing operations such as... Figure 1 The computer program (including program code) involved in each step of the brand keyword acquisition method shown is used to construct, for example... Figure 7 The diagram illustrates a brand term acquisition device and a method for implementing the brand term acquisition method of the embodiments of this application. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the same medium, and executed therein.

[0256] Based on the description of the method and apparatus embodiments above, this application also provides an electronic device. Please refer to... Figure 8The electronic device includes at least a processor 801, an input device 802, an output device 803, and a memory 804. The processor 801, input device 802, output device 803, and memory 804 within the electronic device can be connected via a bus or other means.

[0257] The memory 804 can be stored in the memory of the electronic device. The memory 804 is used to store computer programs, which include program instructions. The processor 801 is used to execute the program instructions stored in the memory 804. The processor 801 (or CPU (Central Processing Unit)) is the computing and control core of the electronic device. It is suitable for implementing one or more instructions, specifically for loading and executing one or more instructions to achieve the corresponding method flow or corresponding function.

[0258] In one embodiment, the processor 801 of the electronic device provided in this application can be used to perform a series of brand term acquisition methods, specifically to perform the following steps:

[0259] Obtain note information, wherein the note information includes note images and note text;

[0260] Subject recognition is performed on the note image to obtain h subjects and h categories, with each of the h subjects and h categories corresponding one-to-one;

[0261] Text detection is performed on the notes to obtain a first set of brand keywords;

[0262] Brand word detection is performed on the notes image to obtain at least one second brand word set;

[0263] Based on the h categories of the h subjects, the first brand word set and the at least one second brand word set are merged to obtain z target brand words.

[0264] Optionally, the electronic device includes, but is not limited to, a processor 801, an input device 802, an output device 803, and a memory 804. It may also include memory, a power supply, an application client module, etc. The input device 802 may be a scanning device, a keyboard, a touchscreen, an RF receiver, etc., and the output device 803 may be a speaker, a display, an RF transmitter, etc. Those skilled in the art will understand that the schematic diagram is merely an example of an electronic device and does not constitute a limitation on the electronic device; it may include more or fewer components than illustrated, or combine certain components, or use different components.

[0265] Optionally, the electronic devices in this application may include smartphones (such as Android phones, iOS phones, Windows Phones, etc.), tablet computers, PDAs, laptops, mobile internet devices (MIDs), or wearable devices. The above-mentioned electronic devices are merely examples and not exhaustive, and include, but are not limited to, the electronic devices described above. In practical applications, the above-mentioned electronic devices may also include: intelligent in-vehicle terminals, computer equipment, etc.

[0266] It should be noted that since the processor 801 of the electronic device executes the steps in the above-described brand term acquisition method when executing the computer program, all embodiments of the above-described brand term acquisition method are applicable to the electronic device and can achieve the same or similar beneficial effects.

[0267] This application embodiment also provides a computer storage medium (Memory), which is a memory device in an electronic device used to store programs and data. It is understood that the computer storage medium here can include the built-in storage medium in a terminal, or it can include an extended storage medium supported by the terminal. The computer storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by the processor 801. These instructions can be one or more computer programs (including program code). It should be noted that the computer storage medium here can be high-speed RAM, or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer storage medium located remotely from the aforementioned processor 801. In one embodiment, the processor 801 can load and execute one or more instructions stored in the computer storage medium to implement the corresponding steps of the aforementioned brand name acquisition method.

[0268] For example, a computer program on a computer storage medium includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0269] It should be noted that since the computer program on the computer storage medium is executed by the processor to implement the steps in the above-described brand name acquisition method, all embodiments of the above-described brand name acquisition method are applicable to the computer storage medium and can achieve the same or similar beneficial effects.

[0270] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0271] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0272] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0273] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0274] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software program module.

[0275] If the integrated unit is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0276] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0277] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for obtaining brand keywords, characterized in that, The method includes: Obtain note information, wherein the note information includes note images and note text; Subject recognition is performed on the note image to obtain h subjects and h categories, with each of the h subjects and h categories corresponding one-to-one; Text detection is performed on the notes to obtain a first set of brand keywords; Brand word detection is performed on the notes image to obtain at least one second brand word set; Based on the h categories of the h subjects, the first brand word set and the at least one second brand word set are merged to obtain z target brand words; The step of fusing the first brand term set and the at least one second brand term set according to the h categories of the h subjects to obtain z target brand terms includes: Determine m first target brand words from the first brand word set and the at least one second brand word set, wherein each first target brand word exists in at least two brand word sets, the first brand word set and the at least one second brand word set; Remove the m first target brand words from the first brand word set and the at least one second brand word set to obtain a third brand word set corresponding to the first brand word set, and a fourth brand word set corresponding to each second brand word set; Based on the h categories, the third brand word set is filtered to obtain n second target brand words; Based on the h categories, brand words are filtered for each fourth brand word set to obtain k third target brand words corresponding to each fourth brand word set; The m first target brand words, the n second target brand words, and the k third target brand words corresponding to each fourth brand word set are merged to obtain the z target brand words.

2. The method according to claim 1, characterized in that, The process involves filtering brand keywords from the third brand keyword set based on the h categories to obtain n second target brand keywords, including: Obtain x fourth target brand words from the third brand word set, wherein the h categories contain the category corresponding to each fourth target brand word; Remove the x fourth target brand terms from the third brand term set to obtain the fifth brand term set; A weighted processing operation is performed on each brand word in the fifth brand word set according to its corresponding category to obtain the target confidence level of each brand word in the fifth brand word set; Based on the target confidence level of each brand word in the fifth brand word set, y fifth target brand words are obtained; The x fourth target brand words and the y fifth target brand words are merged to obtain the n second target brand words.

3. The method according to claim 2, characterized in that, The process involves filtering brand keywords for each fourth brand keyword set based on the h categories, resulting in k third target brand keywords corresponding to each fourth brand keyword set, including: Obtain q sixth target brand words from each fourth brand word set, wherein the h categories contain the category corresponding to each sixth target brand word; Remove the q sixth target brand words from each fourth brand word set to obtain the sixth brand word set corresponding to each fourth brand word set; A weighted processing operation is performed on each brand word in the sixth brand word set according to its corresponding category to obtain the target confidence level of each brand word in the sixth brand word set; Based on the target confidence level of each brand word in the sixth brand word set, determine r seventh target brand words corresponding to each fourth brand word set; The q sixth target brand words in each fourth brand word set and the r seventh target brand words corresponding to each fourth brand word set are merged to obtain the k third target brand words corresponding to each fourth brand word set.

4. The method according to claim 3, characterized in that, The weighting operation includes: Based on the category corresponding to brand word A, determine the penalty coefficient corresponding to brand word A, wherein brand word A is any one in the fifth brand word set or any one in each sixth brand word set; Based on the brand word set in which brand word A is located, obtain the preset penalty item corresponding to brand word A; The target confidence level of brand term A is determined based on the penalty coefficient corresponding to brand term A, the preset penalty item, and the confidence level of brand term A.

5. The method according to claim 4, characterized in that, The text detection of the notes yields a first set of brand keywords, including: Entity recognition is performed on the note text to obtain a first candidate brand words; Map each first-candidate brand term to a main brand term to obtain the brand term corresponding to each first-candidate brand term. The brand words corresponding to each first candidate brand word are combined into the first brand word set.

6. The method according to claim 4, characterized in that, The process of performing brand keyword detection on the note image yields at least one second set of brand keywords, including: Brand logo detection was performed on the note image to obtain b second candidate brand words; Map each second candidate brand term to the main brand term to obtain the brand term corresponding to each second candidate brand term. The brand words corresponding to each second candidate brand word are combined into at least one set of second brand words.

7. The method according to claim 4, characterized in that, The process of performing brand keyword detection on the note image yields at least one second set of brand keywords, including: Optical character recognition was performed on the note image to obtain c third candidate brand words; For each third candidate brand term, perform brand term mapping to obtain the brand term corresponding to each third candidate brand term; The brand words corresponding to each third candidate brand word are combined into at least one set of second brand words.

8. The method according to claim 4, characterized in that, The process of performing brand keyword detection on the note image yields at least one second set of brand keywords, including: Logo detection was performed on the note image to obtain b second candidate brand words; Map each second candidate brand term to the main brand term to obtain the brand term corresponding to each second candidate brand term. Optical character recognition was performed on the note image to obtain c third candidate brand words; Map each second candidate brand term to a main brand term to obtain the brand term corresponding to each third candidate brand term. Each second candidate brand word is combined into a second brand word set, and each third candidate brand word is combined into another second brand word set to obtain the at least one second brand word set.

9. The method according to claim 8, characterized in that, The optical character recognition (OCR) of the note image yields c third candidate brand words, including: Optical character recognition is performed on the note image to obtain text information; The text information is matched with the brand terminology database to obtain d fourth candidate brand terms; The text information is matched with the h categories to obtain e fifth candidate brand words; The d fourth candidate brand words and the e fifth candidate brand words are used as the c third candidate brand words.

10. The method according to claim 9, characterized in that, The process of performing subject recognition on the note image yields h subjects and h categories, including: Object detection is performed on the note image to obtain multiple candidate boxes; Based on the position information of each candidate box, subject detection is performed on the multiple candidate boxes to obtain h subject boxes, wherein the target selected by each subject box is a subject; The subject corresponding to each subject frame is classified to obtain h categories corresponding to each subject.

11. The method according to claim 10, characterized in that, The note information also includes note categories; The note category is used to provide prior information when performing text detection on the note text, so that the brand words in the obtained first brand word set can be matched with the note category; The note category is also used to provide prior information when performing brand word detection on the note image, so that the brand words in each second brand word set obtained are matched with the note category; The note category is also used to provide prior information when performing subject recognition on the note image, so that each subject obtained, and the category of each subject, matches the note category.

12. The method according to claim 11, characterized in that, The method further includes: Based on the z target brand keywords, brand recommendations are made to the user.

13. The method according to any one of claims 1-12, characterized in that, The method further includes: Based on the z target brand keywords, a user profile is constructed for the user.

14. A brand term acquisition device, the device being used to perform the method as described in any one of claims 1-13, characterized in that, include: An acquisition unit is used to acquire note information, wherein the note information includes note images and note text; The processing unit is used to perform subject recognition on the note image to obtain h categories of h subjects, wherein the h subjects and the h categories correspond one-to-one; Text detection is performed on the notes to obtain a first set of brand keywords; Brand word detection is performed on the notes image to obtain at least one second brand word set; Based on the h categories of the h subjects, the first brand word set and the at least one second brand word set are merged to obtain z target brand words.

15. An electronic device, characterized in that, include: A processor and a memory, the processor being connected to the memory, the memory being used to store a computer program, the processor being used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Brand word recognition method and device, equipment and storage medium

    CN110750985A

  • Article brand party determining method and device and server

    CN113836916A