Label sample acquisition method, electronic equipment and storage medium
By crawling the page screenshot of the target page and using the screen understanding model to obtain the description of the topic object and the location of the multimedia file, the problem of obtaining multimedia data annotation samples on different Internet pages is solved, and the automatic acquisition of multimedia data annotation samples on various types of pages is realized.
Patent Information
- Application Number
- CN202510058724.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
It is difficult for the prior art to effectively obtain multimedia data labeling samples of different Internet pages, especially because the Internet pages are ever-changing, it is difficult to set the labeling relationship between pictures and text with simple rules.
By crawling the page screenshot of the target page, input the screenshot into the screen understanding model, obtain the description of the topic object in the target page and the location of the multimedia file, and then obtain and label the multimedia file to obtain the multimedia data labeling sample.
It provides a common multimedia data annotation sample acquisition solution for different types of target pages, which can automatically obtain multimedia data annotation samples, simplifying the crawling process for different pages.
Smart Images

Figure CN119942209A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method for obtaining a labeled sample, an electronic device, and a storage medium. Background Art
[0002] The training of large artificial intelligence models cannot be separated from massive amounts of high-quality data. In related technologies, web crawlers can be used to crawl text and multimedia data such as images, videos, and audio from Internet pages, and the multimedia data can be annotated with the text corresponding to the multimedia data to obtain annotated samples of the multimedia data. However, the pages on the Internet are ever-changing, and it is impossible to use a simple rule to set which picture is annotated with which text. The same is true for multimedia data such as video and audio. How to crawl different pages to obtain annotated samples of multimedia data is a technical problem that needs to be solved urgently. Summary of the invention
[0003] The embodiments of the present application provide a method for obtaining annotated samples, an electronic device, and a storage medium, which are used to provide a universal solution for obtaining annotated samples of multimedia data for different Internet pages.
[0004] In a first aspect, an embodiment of the present application provides a method for obtaining a labeled sample, the method comprising: Crawl a screenshot of the target page; Inputting the page screenshot into a screen understanding model, so that the screen understanding model outputs a description of the subject object in the target page and a location of a multimedia file corresponding to the subject object based on the page screenshot; Acquire the multimedia file corresponding to the subject object based on the multimedia file location; Based on the description of the subject object, the multimedia file corresponding to the subject object is annotated to obtain a multimedia data annotated sample corresponding to the subject object.
[0005] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory having a computer program stored thereon; A processor is used to execute the computer program to implement the method as described in the first aspect of the embodiment of the present application.
[0006] In a third aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method described in the first aspect of the embodiment of the present application.
[0007] In an embodiment of the present application, a page screenshot of a target page is crawled and the page screenshot is input into a screen understanding model, so that the screen understanding model outputs a description of the subject object in the target page and a multimedia file location corresponding to the subject object based on the page screenshot. Then, the multimedia file corresponding to the subject object is obtained based on the multimedia file location, and the multimedia file corresponding to the subject object is annotated based on the description of the subject object to obtain a multimedia data annotation sample corresponding to the subject object. This can provide a general multimedia data annotation sample acquisition scheme for various types of target pages, and facilitate obtaining multimedia data annotation samples based on various types of target pages. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0009] Figure 1 It is a flowchart of a method for obtaining a labeled sample provided in an embodiment of the present application; Figure 2 This is a schematic diagram of a pop-up window processing flow provided in an embodiment of the present application; Figure 3 This is an application example flow chart provided by an embodiment of the present application; Figure 4 It is a module schematic diagram of a device for obtaining a labeled sample provided in an embodiment of the present application; Figure 5 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0010] Embodiments of the present application provide a method for acquiring a labeled sample, an electronic device, and a storage medium.
[0011] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0012] The method for obtaining annotated samples provided in the embodiment of the present application can be executed by an automated tool to achieve automated acquisition of annotated samples of multimedia data. Figure 1 FIG. 1 is a schematic diagram of a flow chart of obtaining a labeled sample provided in an embodiment of the present application. Figure 1 As shown, a method for obtaining a labeled sample provided in an embodiment of the present application includes the following steps: S102, crawling a screenshot of the target page.
[0013] In an embodiment of the present application, a general method for obtaining multimedia data annotation samples is provided for different target pages. There may be multiple target pages. The types of the multiple target pages may be different, and their page layouts, organization methods, etc. may be different. For each target page of multiple different types of target pages, the corresponding multimedia data annotation samples can be obtained based on the method provided in the embodiment of the present application.
[0014] In the embodiment of the present application, the target page is a page based on which multimedia data annotation samples are to be obtained, which changes with the application scenario.
[0015] In one application scenario, the target page can be a page in a user-specified page range. In specific implementation, the URL of the target page specified by the user can be obtained, and a web crawler tool can be called to crawl a page screenshot of the target page based on the URL of the target page.
[0016] In one application scenario, the target page may be a page related to the target object. In a specific implementation, the target object-related pages may be searched in a search engine, and the target page may be screened based on the page screening rules corresponding to the target object to determine the target page. Afterwards, the URL of the target page may be input into a web crawler tool, so that the web crawler tool crawls a page screenshot of the target page based on the URL of the target page.
[0017] The target object may be an object for which annotated samples of multimedia data such as images, videos or audios need to be obtained, such as an item, a product, a plant, an animal, a song, etc. The page filtering rules corresponding to the target object vary with the target object.
[0018] For example, when the target object is a product, the page filtering rule corresponding to the target object can be: filtering out pages that specifically introduce this product and / or specifically sell this product, and the target page can be: the filtered pages that specifically introduce this product (i.e., the introduction page of this product) and / or the page that specifically sells this product (i.e., the sales page of this product).
[0019] When the target object is a plant, the page filtering rules corresponding to the target object can be: filtering out pages that specifically introduce the plant and / or specifically sell the plant, and the target page can be: the filtered pages that specifically introduce the plant (i.e., the introduction page of the plant) and / or the page that specifically sells the plant item (i.e., the product sales page of the plant).
[0020] When the target object is a song, the page filtering rule corresponding to the target object may be: filtering out pages that specifically play the song, and the target page may be: a page that specifically plays the song (ie, the play page of the song).
[0021] S104, inputting the page screenshot into a screen understanding model, so that the screen understanding model outputs a description of the subject object in the target page and a multimedia file location corresponding to the subject object based on the page screenshot.
[0022] In the embodiment of the present application, the subject object in the target page is the core or theme of the content of the target page, which varies with different target pages. For example, when the target page is an introduction page of a certain product, the subject object of the target page can be the product. When the target page is a sales page of a certain plant, the subject object of the target page can be the plant. When the target page is a play page of a certain song, the subject object of the target page can be the song.
[0023] In the embodiment of the present application, the description of the subject object in the target page is a text description of the subject object in the target page, which may include: the name of the subject object, key information, etc. For example, when the target page is an introduction page of a certain product, the description of the subject object in the target page may include: the name of the product and product performance information, etc. When the target page is a sales page of a certain plant, the description of the subject object in the target page may include: the name of the plant, price information, etc. When the target page is a playback page of a certain song, the description of the subject object in the target page may include: the name of the song, the singer, etc.
[0024] In the embodiment of the present application, the multimedia file corresponding to the subject object in the target page can be an image, video, audio, etc. of the subject object in the target page. For example, when the target page is a product introduction page of a certain product, the multimedia file corresponding to the subject object in the target page can be an introduction video of the product in the target page. When the target page is a sales page of a certain plant, the multimedia file corresponding to the subject object in the target page can be a display image of the plant in the target page. When the target page is a playback page of a certain song, the multimedia file corresponding to the subject object in the target page can be the playback audio of the song.
[0025] In the embodiment of the present application, the screen understanding ability of the screen understanding model is used to enable the screen understanding model to understand the target page based on the page screenshot of the target page, and output the description of the subject object in the target page and the multimedia file location corresponding to the subject object. Among them, the screen understanding model is a large language model that can understand the screen content, which has the ability to identify and understand the content in the screen and inform the location of each part of the content. The screen understanding model is, for example, a multimodal large language model (Multimodal Large Language Models, referred to as MLLMs) that can understand the screen content. By inputting the page screenshot of the target page into the screen understanding model, the description of the subject object in the target page and the multimedia file location corresponding to the subject object are obtained by using the screen understanding model, a general subject object description and multimedia file location acquisition method can be provided for different types of target pages, thereby providing a general multimedia data annotation sample acquisition method for different types of target pages. Compared with the page layout of each type of target page, etc., a corresponding crawling method is specially designed for each type of target page to obtain its subject object description and the multimedia file corresponding to the subject object, and obtain the multimedia data annotation sample corresponding to the subject object, the method provided in the embodiment of the present application is simple to implement and can be applied to various types of target pages.
[0026] In a specific implementation, appropriate prompt words can be used to guide the screen understanding model to identify and output the description of the subject object in the target page and the location of the multimedia file corresponding to the subject object based on the page screenshot of the target page.
[0027] For example, when the target page is a product introduction page, a first prompt word may be generated, and the first prompt word is used to guide the screen understanding model to output the product name of the product introduced in the introduction page based on the page screenshot of the introduction page, which is, for example, "This is a screenshot of a product introduction page. What product does this page introduce?" Then, the first prompt word and the page screenshot of the introduction page may be input into the screen understanding model, and the screen understanding model may output the product name of the product introduced in the introduction page under the guidance of the first prompt word. After that, a second prompt word may be generated, and the second prompt word is used to guide the screen understanding model to output the location of the product image based on the product introduction screenshot, which is, for example, "Which images are related to the introduced product in this page?" Then, the second prompt word may be input into the screen understanding model, and the screen understanding model may mark the location of the product picture and / or video in the page screenshot under the guidance of the second prompt word.
[0028] For another example, when the target page is a sales page of a product, a first prompt word may be generated, and the first prompt word is used to guide the screen understanding model to output the product name of the product sold in the sales page based on the page screenshot of the sales page, which is, for example, "This is a screenshot of a product sales page. What product is sold on this page?" Then, the first prompt word and the page screenshot of the sales page may be input into the screen understanding model, and the screen understanding model may output the name of the product sold in the product sales page under the guidance of the first prompt word. After that, a second prompt word may be generated, and the second prompt word is used to guide the screen understanding model to output the image location of the product sold based on the product sales screenshot, which is, for example, "Which images are related to the product sold on this page?" Then, the second prompt word may be input into the screen understanding model, and the screen understanding model may mark the location of the picture and / or video of the product in the item sales screenshot under the guidance of the second prompt word.
[0029] For another example, when the target page is a play page of a song, a first prompt word may be generated, and the first prompt word is used to guide the screen understanding model to output the song name of the song played on the play page based on the page screenshot of the play page, which is, for example, "This is a screenshot of the play page of a song. What song is played on this play page?" Then, the first prompt word and the screenshot of the song play page may be input into the screen understanding model, and the screen understanding model may output the name of the song played on the song play page under the guidance of the first prompt word. After that, a second prompt word may be generated, and the second prompt word is used to guide the screen understanding model to output the audio position of the song played based on the screenshot of the song play page, which is, for example, "In this page, which are the audios related to the song played?" Then, the second prompt word may be input into the screen understanding model, and the screen understanding model may mark the playback audio position of the song in the screenshot of the song play page under the guidance of the second prompt word.
[0030] S106: Acquire the multimedia file corresponding to the subject object based on the multimedia file position.
[0031] In an embodiment of the present application, the multimedia file location is the location marked by the screen understanding model in the page screenshot of the target page, and the original file of the multimedia file corresponding to the subject object can be obtained based on this location.
[0032] Specifically, the screen coordinates of the multimedia file can be calculated based on the multimedia file location; based on the screen coordinates, the original URL (Uniform Resource Locator) of the multimedia file can be looked up in the page source code of the target page; and the original file of the multimedia file can be downloaded based on the original URL.
[0033] For example, when the target page is an introduction page of a product, the subject object in the target page may be the product, and the corresponding multimedia file location of the subject object in the target page may be the location of the introduction picture / introduction video of the product in the target page. Based on the location of the introduction picture / introduction video, the screen coordinates of the introduction picture / introduction video may be calculated, and based on the screen coordinates, the original URL of the introduction picture / introduction video may be reversely checked in the page source code of the target page; then, the original file of the introduction picture / introduction video may be downloaded based on the original URL.
[0034] For another example, when the target page is a sales page for a product, the subject object in the target page may be the product, and the corresponding multimedia file location of the subject object in the target page may be the location of the display image / display video of the product in the target page. Based on the location of the display image / display video, the screen coordinates of the display image / display video may be calculated, and based on the screen coordinates, the original URL of the display image / display video may be reversely checked in the page source code of the target page; then, the original file of the display image / display video may be downloaded based on the original URL.
[0035] For another example, when the target page is a song playing page of a song, the subject object in the target page can be the song, and the corresponding multimedia file position of the subject object in the target page can be the position of the playing audio of the song in the target page. Based on the position of the playing audio of the song, the screen coordinates of the playing audio of the song can be calculated, and based on the screen coordinates, the original URL of the playing audio can be reversed in the page source code of the target page; then, the original file of the playing audio of the song can be downloaded based on the original URL.
[0036] In a specific implementation, the screen understanding model can frame the location of the multimedia file in the page screenshot of the target page, and can determine the screen location of the multimedia file based on the center point of the multimedia file location framed by the screen understanding model, and then, based on the screen location, reversely search the original URL of the multimedia file from the source code of the target page. When the original URL of the multimedia file cannot be reversely searched based on the screen location, the screen location can be offset, and the original URL of the multimedia file can be reversely searched again based on the offset screen location until the original URL is reversely searched.
[0037] S108, annotating multimedia files corresponding to the subject object based on the description of the subject object to obtain annotated multimedia data samples corresponding to the subject object.
[0038] Specifically, the original file of the multimedia file corresponding to the subject object may be annotated based on the description of the subject object to obtain an annotated sample of the multimedia data corresponding to the subject object.
[0039] For example, if the subject object is a product, the description of the subject object is the product name of the product, and the multimedia file corresponding to the subject object is the introduction video of the product, the product name of the product can be used to annotate the original video of the introduction video of the product to obtain a video annotation sample of the product.
[0040] If the subject object is a product, the description of the subject object is the product name of the product, and the multimedia file corresponding to the subject object is the display image of the product, the product name of the product can be used to annotate the original image of the display image of the product to obtain the image annotation sample of the product.
[0041] If the subject object is a song, the description of the subject object is the song title, and the multimedia file corresponding to the subject object is the playback audio of the song. The song title of the song can be used to annotate the original audio of the playback audio to obtain an audio annotation sample of the song.
[0042] The method for obtaining multimedia data annotation samples provided in the embodiment of the present application crawls a page screenshot of a target page, inputs the page screenshot into a screen understanding model, and enables the screen understanding model to output a description of a subject object in the target page and a multimedia file location corresponding to the subject object based on the page screenshot. Then, the multimedia file corresponding to the subject object is obtained based on the multimedia file location, and the multimedia file corresponding to the subject object is annotated based on the description of the subject object to obtain a multimedia data annotation sample corresponding to the subject object. This method can provide a general method for obtaining multimedia data annotation samples for various types of target pages, and facilitates obtaining multimedia data annotation samples based on various types of target pages. In a specific implementation, in the above step S102, a crawler tool may be specifically called to simulate opening the target page based on the URL of the target page, and to take a screenshot of the target page, thereby obtaining a page screenshot of the target page. In some cases, when the target page is opened, there may be a pop-up window, and the pop-up window image may block the relevant information in the target page, affecting the screen understanding model's understanding of the target page.
[0043] Figure 21 is a schematic diagram of a pop-up window processing flow provided by an embodiment of the present application. In one implementation, in order to avoid the influence of the pop-up window in the target page on the understanding of the target page, before step S104, the method provided by the embodiment of the present application further includes the following steps: S202, inputting the page screenshot into a screen understanding model, so that the screen understanding model determines whether there is a pop-up window in the target page based on the page screenshot.
[0044] Specifically, appropriate prompt words can be used to guide the screen understanding model to determine whether there is a pop-up window in the target page.
[0045] For example, a third prompt word can be generated, and the third prompt word is used to guide the screen understanding model to determine whether there is a pop-up window in the target page based on the page screenshot of the target page. For example, "The following image is a screenshot of a page. Is there a pop-up window in the page?" Then, the third prompt word and the page screenshot of the target page can be input into the screen understanding model. The screen understanding model can output information on whether there is a pop-up window in the target page under the guidance of the third prompt word.
[0046] S204, when there is a pop-up window in the target page, enable the screen understanding model to determine the user operation required to eliminate the pop-up window based on the page screenshot.
[0047] Continuing with the above example, a fourth prompt word can be further generated. The fourth prompt word is used to guide the screen understanding model to further determine the user operation required to eliminate the pop-up window based on the page screenshot. For example, it is: "What operation does the user need to perform to eliminate the pop-up window?" Then, the fourth prompt word can be input into the screen understanding model. The screen understanding model can output information about the user operation required to eliminate the pop-up window under the guidance of the fourth prompt word.
[0048] S206: Perform a simulation operation of the user operation on the pop-up window to eliminate the pop-up window.
[0049] S208: reacquire a page screenshot of the target page.
[0050] Specifically, a crawler tool may be called to simulate the user operation on the pop-up window in the product sales page based on the pop-up window information and user operation information output by the screen understanding model to eliminate the pop-up window, and then, a screenshot of the target page may be taken again to obtain a page screenshot of the target page. After that, the above steps S202 to S208 may be iteratively executed until there is no pop-up window in the target page.
[0051] For example, when the screen understanding model outputs a login pop-up window in the target page and the user is required to perform a login operation, the crawler tool can be called to simulate the user operation based on pre-saved login information such as account and password, and fill in the login information in the pop-up window to eliminate the pop-up window. Afterwards, the target page can be screenshoted again, and the screenshot of the target page obtained can be input into the screen understanding model to re-determine whether there is a pop-up window in the target page.
[0052] For another example, when the screen understanding model outputs that there is an advertising pop-up window in the target page, and the user is required to click a close button or an exit button in the pop-up window, the crawler tool can be called to simulate clicking the close button or the exit button in the pop-up window to eliminate the pop-up window. Afterwards, the target page can be re-shot, and the re-obtained target page screenshot can be input into the screen understanding model to re-determine whether there is a pop-up window in the target page.
[0053] Furthermore, in the case where there is no pop-up window in the product sales page, step S104 may be executed.
[0054] By iteratively executing the above steps S202 to S208 before step S104, it is possible to avoid the pop-up window affecting the screen understanding model's understanding of the target page, thereby improving the accuracy of the screen understanding model's understanding of the target page.
[0055] The following is a further explanation of the method provided in the embodiment of the present application by taking the application scenario of obtaining image annotation samples of the target object as an example. Figure 3 This is an example flow chart of an application provided by the embodiment of the present application. Figure 3 As shown, in an application example of an embodiment of the present application, the following process is included: S302, searching for pages related to the target object in a search engine.
[0056] For example, if the target object is the plant "Goosegrass", you can search for "Goosegrass" as a keyword in the search engine to get multiple pages containing the keyword "Goosegrass". These pages are pages related to "Goosegrass".
[0057] S304: Screen pages related to the target page based on the page screening rule corresponding to the target object to determine the target page.
[0058] Continuing with the above example, the page screening rules corresponding to the target object may be: screening out pages that specifically introduce "beef tendon grass" and pages that specifically sell "beef tendon grass". Multiple target pages may be: among multiple pages containing the keyword "beef tendon grass", the introduction page of "beef tendon grass" and the sales page of "beef tendon grass". The page layout of the introduction page of "beef tendon grass" and the sales page of "beef tendon grass" may be different. For example, the introduction page of "beef tendon grass" may be the entry page of the entry "beef tendon grass" in an online encyclopedia, and the sales page of "beef tendon grass" may be the product sales page of the product "beef tendon grass" on an e-commerce website.
[0059] When multiple target pages are determined, the following steps S306 to S316 are performed for each target page to obtain a corresponding image annotation sample based on each target page.
[0060] S306, crawling a screenshot of the target page.
[0061] Specifically, the URL of the target page can be passed into a web crawler tool, so that the web crawler tool opens the target page based on the URL and takes a screenshot, thereby obtaining a page screenshot of the target page.
[0062] S308: Input the page screenshot into a screen understanding model, so that the screen understanding model determines whether there is a pop-up window in the target page based on the page screenshot.
[0063] Among them, when there is a pop-up window in the target page, the screen understanding model determines the user operation required to eliminate the pop-up window based on the page screenshot, simulates the user operation on the pop-up window to eliminate the pop-up window, and then re-takes a screenshot of the target page, and re-executes step S308 for the re-obtained page screenshot until there is no pop-up window in the target page.
[0064] S310, when there is no pop-up window in the target page, the screen understanding model is enabled to output the name of the subject object in the target page and the image position corresponding to the subject object based on the page screenshot.
[0065] Continuing with the above example, when the target page is the introduction page of "Beef Tendon Grass", the subject object in the target page is "Beef Tendon Grass". Based on the screen understanding model, it can be determined that the description of the subject object in the target page includes: the name of the plant "Beef Tendon Grass" and the morphological characteristics, growth environment, uses and other information of the plant "Beef Tendon Grass". The image position corresponding to the subject object is the position of the introduction image of the plant "Beef Tendon Grass".
[0066] When the target page is the sales page of "Beef Tendon Grass", the subject object in the target page is "Beef Tendon Grass". Based on the screen understanding model, it can be determined that the description of the subject object in the target page includes information such as the name of the plant "Beef Tendon Grass" and the selling price of the plant "Beef Tendon Grass". The image position corresponding to the subject object is the position of the display image of the plant "Beef Tendon Grass", such as the position of the main picture of the product in the product sales page of "Beef Tendon Grass", which is the position of the display image of "Beef Tendon Grass" in the product sales page of "Beef Tendon Grass".
[0067] S312: Based on the image position corresponding to the subject object in the target page, reversely search the original URL of the image corresponding to the subject object from the page source code of the target page.
[0068] S314: Acquire an image corresponding to the subject object based on the original URL.
[0069] Continuing with the above example, when the target page is the introduction page of "Goosegrass", the subject object of the target webpage is "Goosegrass". Based on the location of the introduction image of the plant "Goosegrass" in the introduction page, the original URL of the introduction image can be reversed from the source code of the introduction page, and then the original image of the introduction image can be obtained from the original URL.
[0070] When the target page is the sales page of "Goosegrass", the subject object of the target webpage is "Goosegrass". Based on the location of the display image of the plant "Goosegrass" in the sales page, the original URL of the display image can be reversed from the source code of the sales page, and then the original image of the display image can be obtained from the original URL.
[0071] The introduction image or the display image may be a static image, ie, a picture, or a dynamic image, ie, a video.
[0072] S316, annotating the image corresponding to the subject object based on the description of the subject object to obtain an annotated image sample corresponding to the subject object.
[0073] Using the above example, when the target page is the introduction page of "Goosegrass", the introduction image of "Goosegrass" can be annotated based on the name, morphological characteristics, growth environment and other descriptions of "Goosegrass" in the introduction page to obtain the image annotation sample of "Goosegrass". When the target web page is the sales page of "Goosegrass", the display image of "Goosegrass" can be annotated based on the name, price and other descriptions of "Goosegrass" in the sales page to obtain the image annotation sample of "Goosegrass".
[0074] Among them, in the above example, when the target web page is the sales page of the target object, in the above steps S312~S314, the main image of the product of the subject object can be obtained based on the position of the main image of the product of the subject object in the target web page. Thereafter, in step S316, the description of the subject object in the target web page can be used to annotate the main image of the product of the subject object, thereby obtaining an image annotation sample corresponding to the subject object.
[0075] Furthermore, in the case where the target web page is a sales page for the target object, after step S316, the main image of the product of the subject object can be further switched based on the product carousel of the subject object in the target web page, and then, based on the position of the main image of the product of the subject object in the target web page, the switched main image of the product is obtained, and the switched main image of the product is annotated using descriptions such as the name of the subject object in the target web page, thereby obtaining another image annotation sample of the subject object.
[0076] Specifically, a page screenshot of the target page can be input into a screen understanding model, so that the screen understanding model outputs the position of the product carousel of the subject object in the target page based on the page screenshot. Then, based on the position of the product carousel, a simulated click operation can be performed on the product carousel in the product sales page to switch the main image of the product at the position of the product main image. After that, the switched main image of the product can be obtained based on the position of the product main image, and the switched main image of the product can be labeled based on the description of the subject object, thereby obtaining an image annotation sample corresponding to the subject object.
[0077] The main product image can be used to display the main image of the target object in the sales page of the target object, and the product carousel image is an image in the sales page that is clicked to switch the main product image. For example, when the target object is the plant "Goosegrass", the subject object in the target webpage is also "Goosegrass", the main product image can be the main display image of "Goosegrass" in the sales page of the Goosegrass, such as a large image of "Goosegrass", and the product carousel image can be a thumbnail display image of "Goosegrass" in the sales page of the product.
[0078] By switching the main product image of the subject object in the sales page based on the product carousel image in the sales page when the target page is the sales page of the target object, and annotating the switched main product image, multiple image annotation samples corresponding to the subject object of the page can be obtained based on the same page, thereby further improving the diversity and richness of the annotation samples.
[0079] The above is a method for obtaining labeled samples provided in an embodiment of the present application. Based on the same idea, an embodiment of the present application also provides a device for obtaining labeled samples. Figure 4Schematic diagram of a module of a device for acquiring multimedia data annotation samples provided in an embodiment of the present application. Figure 4 As shown, the device for obtaining multimedia data annotated samples provided in the embodiment of the present application includes the following modules: A screenshot crawling module 401 is used to crawl page screenshots of a target page; A page understanding module 402, configured to input the page screenshot into a screen understanding model, so that the screen understanding model outputs a description of a subject object in the target page and a location of a multimedia file corresponding to the subject object based on the page screenshot; A file acquisition module 403, used to acquire the multimedia file corresponding to the subject object based on the multimedia file position; The labeling module 404 labels the multimedia files corresponding to the subject object based on the description of the subject object to obtain a labeled sample of multimedia data corresponding to the subject object.
[0080] The device for acquiring labeled samples provided in the embodiment of the present application and the method for acquiring labeled samples provided in the above embodiment are based on the same inventive concept and can achieve the same technical effect, which will not be described in detail.
[0081] The methods and devices provided in the embodiments of the present application can be applied to electronic devices. Figure 5 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device includes a processor 501 and a memory 502. The memory 502 stores a computer program that can be run on the processor 501. When the computer program is executed by the processor 501, the various processes of the above-mentioned labeled sample acquisition method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0082] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned method for obtaining annotated samples is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0083] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0084] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0085] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0087] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0088] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0089] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. According to the definition in this article, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0090] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0091] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for obtaining a labeled sample, characterized in that: The method comprises: Crawl a screenshot of the target page; Inputting the page screenshot into a screen understanding model, so that the screen understanding model outputs a description of the subject object in the target page and a location of a multimedia file corresponding to the subject object based on the page screenshot; Acquire the multimedia file corresponding to the subject object based on the multimedia file location; Based on the description of the subject object, the multimedia file corresponding to the subject object is annotated to obtain a multimedia data annotated sample corresponding to the subject object.
2. The method according to claim 1, characterized in that The multimedia files include: pictures, videos or audios.
3. The method according to claim 1, characterized in that The acquiring the multimedia file corresponding to the subject object based on the multimedia file position includes: Calculating the screen coordinates of the multimedia file based on the multimedia file position; Based on the screen coordinates, searching the page source code of the target page for the original URL of the multimedia file; Downloading the original file of the multimedia file based on the original URL; The step of labeling the multimedia file corresponding to the subject object based on the description of the subject object to obtain a multimedia data labeling sample corresponding to the subject object includes: The original file of the multimedia file is annotated based on the description of the subject object to obtain a multimedia data annotated sample corresponding to the subject object.
4. The method according to claim 1, characterized in that: Before inputting the page screenshot into a screen understanding model so that the screen understanding model outputs a description of the subject object in the target page and a location of a multimedia file corresponding to the subject object based on the page screenshot, the method further includes: Inputting the page screenshot into a screen understanding model, so that the screen understanding model determines whether there is a pop-up window in the target page based on the page screenshot; When there is a pop-up window in the target page, the screen understanding model determines the user operation required to eliminate the pop-up window based on the page screenshot; Performing a simulated operation of the user operation on the pop-up window to eliminate the pop-up window; Re-acquire a page screenshot of the target page.
5. The method according to claim 1, characterized in that Before crawling the page screenshot of the target page, the method further includes: Search for pages related to the target object in search engines; Based on the page screening rule corresponding to the target object, pages related to the target object are screened to determine the target page.
6. The method according to claim 5, characterized in that The target page includes: an introduction page of the target object, and / or a sales page of the target object.
7. The method according to claim 6, characterized in that In the case where the target page is a sales page of the target object, the multimedia file location includes: a location of a main image of a product corresponding to the subject object in the target page, and the multimedia file corresponding to the subject object is annotated based on the description of the subject object to obtain a multimedia data annotation sample corresponding to the subject object, including: Based on the description of the subject object, the main image of the commodity of the subject object is annotated to obtain an image annotation sample corresponding to the subject object.
8. The method according to claim 7, characterized in that After obtaining the image annotation sample of the target object, the method further includes: Inputting the page screenshot into a screen understanding model, so that the screen understanding model outputs the position of the product carousel image of the subject object in the target page based on the page screenshot; Based on the position of the product carousel image, a simulated click operation is performed on the product carousel image in the sales page to switch the product main image at the position of the product main image; Obtaining a switched main image of the product based on the position of the main image of the product; The switched main image of the product is annotated based on the description of the subject object to obtain an image annotation sample corresponding to the subject object.
9. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the method as claimed in any one of claims 1 to 8.