Information pushing method and device, terminal equipment and storage medium
By acquiring the corresponding image set of graffiti images and pushing target content based on the category information selected by the user, the problem of the single graffiti image recognition result of terminal devices is solved, and the diversity and accuracy of the recognition results are improved.
Patent Information
- Application Number
- CN202210408927.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-04-19
AI Technical Summary
In existing technologies, the recognition results of graffiti images by terminal devices are relatively simple, with low recognition accuracy and flexibility, and users cannot flexibly select and understand the relevant content of graffiti images.
The system retrieves the image set corresponding to the graffiti image from the terminal device, responds to the user's selection of the first image, obtains its category information, and pushes target content based on the category information to describe the target object in the image.
It improves the diversity and accuracy of graffiti image recognition results, allowing users to easily understand the content of the drawn graffiti images, and enhances the flexibility and accuracy of recognition.
Smart Images

Figure CN116975348B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an information push method, apparatus, terminal device and storage medium. Background Technology
[0002] With the rapid development of science and technology, there are more and more scenarios in daily life where terminal devices are used. For example, people use terminal devices to take pictures or to transmit messages.
[0003] For functions like drawing and doodling, people can use terminal devices to teach children. After installing an application (app) with drawing capabilities on the terminal device, the user can identify the drawn image and obtain a set of corresponding image resources. For example, after a child doodles in a drawing app, parents can use their own understanding to identify the doodle and explain it to the child, providing more guidance. Currently, in the process of recognizing doodles using terminal devices, it is often based on trained machine learning models. This process tends to produce relatively simple and low-accuracy results. Summary of the Invention
[0004] This application provides an information push method, apparatus, terminal device, and storage medium, which can obtain category information from the recognition results of graffiti images through human-computer interaction, and push corresponding target content based on the category information, so that users can more conveniently understand the recognition results of graffiti images and improve the diversity and accuracy of recognition results.
[0005] In one aspect, embodiments of this application provide an information push method, the method being applied to a terminal device, the method comprising:
[0006] Get the graffiti image;
[0007] Based on the graffiti image, obtain the corresponding set of source images;
[0008] In response to the selection operation of the first image, the category information of the first image is obtained. The first image is any one of the material images in the set. The category information is used to indicate the category of the target object in the first image.
[0009] Based on the category information, target content is pushed to the first image, whereby the target content describes the target object of the first image.
[0010] Optionally, the terminal device includes a display screen displaying a drawing interface, the drawing interface including functional areas; the step of obtaining category information of the first image in response to a selection operation on the first image includes:
[0011] In response to a touch operation on a selection control in the functional area, the first image is acquired;
[0012] Based on the first image, obtain the category information of the first image.
[0013] Optionally, the graffiti interface further includes a push area, wherein pushing the target content of the first image according to the category information includes:
[0014] Based on the category information, obtain the target content of the first image;
[0015] The target content is pushed to the push area according to a preset push method.
[0016] Optionally, the terminal device includes a display screen, and before acquiring the graffiti image, it further includes:
[0017] The drawing interface is displayed on the screen, and the drawing interface also includes a drawing area and a tool area;
[0018] A doodle image is generated in response to a drawing operation in the drawing area;
[0019] The process of obtaining the graffiti image includes:
[0020] Obtain the graffiti image generated in the graffiti interface; or, in response to a trigger operation on the camera control in the functional area, obtain the imported graffiti image.
[0021] Optionally, obtaining the set of source images corresponding to the graffiti image includes:
[0022] The graffiti image is input into the image generation model to generate a first set of filtered images;
[0023] Based on the image parameters of each image in the first filtered image set, obtain the similarity between each image in the first filtered image set and the graffiti image;
[0024] Based on the similarity of each image in the first filtered image set, images with a similarity higher than a preset similarity threshold are selected from the first filtered image set as the material image set.
[0025] Optionally, after inputting the graffiti image into the image generation model to generate the first set of filtered images, the method further includes:
[0026] Get the number of images in the first filtered image set;
[0027] When the number of images in the first filtered image set is not greater than a preset threshold, the graffiti image is input into the image search model to obtain the second filtered image set. Each image obtained by the image search model contains its own confidence level.
[0028] Based on the confidence level of each image in the second filtered image set, the material image set is obtained from the second filtered image set;
[0029] When the number of images to be filtered exceeds the preset threshold, the step of obtaining the similarity between the images to be filtered and the graffiti images based on the image parameters of the images to be filtered is executed.
[0030] Optionally, after pushing the target content of the first image according to the category information, the method further includes:
[0031] In response to a trigger operation on the target content, a content explanation interface is displayed;
[0032] The target content of the first image will be explained according to the preset explanation method.
[0033] In another aspect, embodiments of this application provide an information push device, which is applied to a terminal device, and the device includes:
[0034] The first acquisition module is used to acquire graffiti images;
[0035] The second acquisition module is used to acquire the material image set corresponding to the graffiti image based on the graffiti image;
[0036] The third acquisition module is used to acquire category information of the first image in response to the selection operation of the first image, wherein the first image is any one of the material images in the set, and the category information is used to indicate the category of the target object in the first image;
[0037] The information push module is used to push target content of the first image according to the category information, wherein the target content is used to describe the target object of the first image.
[0038] In another aspect, embodiments of this application provide a terminal device, which includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor enables the processor to implement the information push method as described in one aspect above.
[0039] In another aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the information push method as described in the other aspect and its alternatives above.
[0040] On the other hand, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to execute the information push method as described in one aspect above.
[0041] On the other hand, embodiments of this application provide an application publishing platform for publishing computer program products, wherein when the computer program product is run on a computer, the computer executes the information push method as described in one aspect above.
[0042] The technical solutions provided in this application embodiment may include at least the following beneficial effects:
[0043] This application acquires graffiti images via a terminal device; based on the graffiti images, it acquires a set of corresponding source images; in response to a selection operation on a first image, it acquires the category information of the first image, where the first image is any one of the source images, and the category information indicates the category of the target object in the first image; based on the category information, it pushes the target content of the first image, where the target content describes the target object of the first image. This application acquires a set of source images from graffiti images, allows selection of one image through human-computer interaction, acquires the category information of the selected image, and acquires and pushes the target content of the selected image based on this category information. This makes it easier to understand the graffiti images drawn by the user and improves the diversity and accuracy of the recognition results. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a schematic diagram of an architecture for a user-drawn doodle image according to an exemplary embodiment of this application;
[0046] Figure 2 This is a flowchart of an exemplary embodiment of the present application for an information push method;
[0047] Figure 3 This is a flowchart of an exemplary embodiment of the present application for an information push method;
[0048] Figure 4 This is a schematic diagram of a graffiti interface according to an exemplary embodiment of this application;
[0049] Figure 5 This is a schematic diagram of a content explanation interface according to an exemplary embodiment of this application;
[0050] Figure 6 This is a schematic diagram of a graffiti interface according to an exemplary embodiment of this application;
[0051] Figure 7 This is a structural block diagram of an information push device provided in an exemplary embodiment of this application;
[0052] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an exemplary embodiment of this application. Detailed Implementation
[0053] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0054] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0055] It should be noted that the terms "first," "second," "third," and "fourth," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "having," and any variations thereof, in the embodiments of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0056] The solution provided in this application can be applied to scenarios where people need to recognize doodles after they have doodled on terminal devices with drawing functions in their daily lives. To facilitate understanding, some terms and application architectures involved in the embodiments of this application will be briefly introduced below.
[0057] Generative Adversarial Networks (GANs) are deep learning models and one of the most promising unsupervised learning methods on complex distributions in recent years. The model produces reasonably good outputs through a game-like learning process between (at least) two modules: a generative model and a discriminative model. In the original GAN theory, it is not required that both G and D be neural networks; they only need to be able to fit the corresponding generative and discriminative functions. However, in practice, deep neural networks are generally used as both G and D.
[0058] With the continuous advancement of technology, terminal devices are being used more and more frequently in daily life. People can use these devices for learning, entertainment, work, and other activities. Among them, terminal devices with drawing apps installed allow users to open the app and draw, creating their own doodles.
[0059] Please refer to Figure 1 This illustrates a schematic diagram of an architecture for user-drawn doodle images according to an exemplary embodiment of this application. Figure 1 As shown, it includes terminal device 110 and server 120.
[0060] Optionally, the terminal device is a device with a drawing function app installed. For example, the terminal device can be a mobile phone, tablet computer, laptop computer, e-book reader, smart glasses, smartwatch, laptop computer, desktop computer, smart robot, smart bracelet, tutoring machine, etc.
[0061] Optionally, the user can draw by opening an app with drawing function on the terminal device and obtain a doodle image displayed on the interface of the terminal device 110. If the user needs to recognize the doodle image, they can send the doodle image to the server through the communication network connection between the terminal device and the server, and the server will recognize the doodle image.
[0062] Optionally, the communication network connection between terminal device 110 and server 120 can use standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to any combination of Local Area Network (LAN), Metropolitan Area Network (MAN), Wide Area Network (WAN), mobile, wired or wireless networks, private networks, or Virtual Private Networks. In some embodiments, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0063] In real-world scenarios, children typically use these drawing functions to create doodles, and both children and parents can understand the content of these doodles through this recognition method. Currently, the methods used for recognition by terminal devices or servers are relatively simplistic. They usually rely on fixed machine learning models to identify doodles, obtain similar images, and then provide these similar images to the user. This process results in overly simplistic information, preventing users from making flexible choices or understanding the relevant content of the images. Consequently, the flexibility and accuracy of recognizing user doodles are both low.
[0064] To enable users to more conveniently understand the recognition results of graffiti images and improve the diversity and accuracy of the recognition results, this application proposes a solution that can obtain images from multiple material image sets based on graffiti images, and push the selected image according to category information based on the user's selection operation of one of the images, thereby improving the efficiency of pushing the actual content of graffiti images.
[0065] Please refer to Figure 2This document illustrates a flowchart of an information push method provided in an exemplary embodiment of this application. This method can be applied to a terminal device and is executed by the terminal device. Figure 2 As shown, this information push method includes the following steps.
[0066] Step 201: Obtain the graffiti image.
[0067] In this application, the terminal device can be equipped with an app that has drawing functions. The user can open the app on the terminal device and draw to obtain the drawn doodle image.
[0068] Step 202: Based on the graffiti image, obtain the corresponding set of source images.
[0069] Optionally, the terminal device obtains a set of source images corresponding to the acquired graffiti image. The images in the source image set can be obtained by the terminal device based on a pre-set machine learning model. The machine learning model involved in this application is obtained through pre-training. The input of the machine learning model is a graffiti image, and the output of the machine learning model is images from multiple source image sets corresponding to the graffiti image. In this step, the source image set is obtained based on the graffiti image.
[0070] Step 203: In response to the selection operation of the first image, obtain the category information of the first image. The first image is any one of the material images in the set, and the category information is used to indicate the category of the target object in the first image.
[0071] Optionally, the terminal device can provide a selection function for the set of source images, allowing the user to choose any image from it. In response to the user's selection of the first image, the terminal device obtains the category information of that first image. For example, after displaying a set of source images in the interface for drawing doodles, the terminal device provides a selection control for the set of source images. The user can use this selection control to choose any image from the set. After the user selects an image from the set, the terminal device obtains the corresponding category information based on that image.
[0072] Optionally, this category information can be obtained by the terminal device during the process of obtaining images from two material image sets based on the graffiti image, or it can be obtained by inputting the selected first image into the category acquisition model in this step, and obtaining the corresponding category information of the first image through the category acquisition model.
[0073] Step 204: Based on the category information, push the target content of the first image. The target content is used to describe the target object of the first image.
[0074] Optionally, the terminal device can push target content to the first image based on the obtained category information. This target content can be informational content corresponding to the target object in the first image. For example, if a user creates a doodle (drawing a horse), step 102 above can obtain at least two images from a set of source images (images of horses in real life). After the user selects one image, the category information for that image is "Red Hare." In this step, based on the "Red Hare" category, information about the Red Hare can be obtained. This information about the Red Hare constitutes the data content corresponding to the target object. The terminal device pushes this target content, thus providing information about the images in the source image set corresponding to the user's doodle, helping the user understand what they drew.
[0075] In summary, this application acquires graffiti images via a terminal device; obtains a set of corresponding source images based on the graffiti images; in response to a selection operation on a first image, it acquires the category information of the first image, where the first image is any one of the source images, and the category information indicates the category of the target object in the first image; and based on the category information, it pushes the target content of the first image, where the target content describes the target object of the first image. This application acquires a set of source images from graffiti images, allows selection of one image through human-computer interaction, acquires the category information of the selected image, and then acquires and pushes the target content of the selected image based on this category information. This makes it easier to understand the graffiti images drawn by the user and improves the diversity and accuracy of the recognition results.
[0076] In one possible implementation, the terminal device can draw on a displayed drawing interface, acquire the drawing image, and push the target content of the first image within the drawing interface. This allows for multiple operations on the drawing image on the same interface, makes it more convenient to push the target content of the first image, and also improves the diversity of the display of the recognition results of the drawing image.
[0077] Please refer to Figure 3 This document illustrates a flowchart of an information push method provided in an exemplary embodiment of this application. This method can be applied to a terminal device and is executed by the terminal device. Figure 3 As shown, this information push method includes the following steps.
[0078] Step 301: Display the drawing interface on the screen. The drawing interface also includes a drawing area and a tool area.
[0079] In this scenario, when a user opens a drawing app on their device, a drawing interface will be displayed on the device's screen. Please refer to [link / reference]. Figure 4This illustration shows a schematic diagram of a graffiti interface according to an exemplary embodiment of this application. Figure 4 As shown, the drawing interface 400 includes a drawing area 401 and a tool area 402. The tool area 402 provides a variety of drawing tools, such as pencil controls, eraser controls, and fill controls. Users can select a drawing tool in the tool area 402 and draw within the drawing area 401.
[0080] Step 302: In response to the drawing operation in the drawing area, generate a doodle image.
[0081] In this graffiti interface, after selecting a drawing tool in the tool area, the user controls that tool to move within the drawing area to create a drawing and generate a graffiti image. Optionally, different drawing tools can correspond to different drawing operations. For example, the pencil tool can be paired with a swipe or a click operation, while the fill tool can be paired with a click operation. In other words, after selecting a drawing tool, the user draws within the drawing area according to their desired style, and the terminal device responds to the drawing operation in the drawing area by generating the corresponding graffiti image.
[0082] Step 303: Obtain the graffiti image.
[0083] Optionally, the terminal device can acquire the graffiti image generated in the graffiti interface; it can also acquire the imported graffiti image in response to a trigger operation of the camera control in the functional area. For example, after the user finishes drawing in the drawing area displayed on the terminal device, the control in the functional area of the graffiti interface that acquires the drawn graffiti image can be triggered, selecting the graffiti image drawn by the user. For example, in the above... Figure 4 The graffiti interface 400 also includes a functional area 403, which contains a "Guess" control 403a. Users can trigger the "Guess" control 403a to allow the terminal device to obtain the graffiti image drawn by the user.
[0084] In one possible implementation, the graffiti interface displayed on the terminal device includes a camera control 403b within the functional area 403. Users can click the camera control 403b to take a picture of the drawn graffiti image using the terminal device, or import graffiti images from their photo album using the terminal device's file import function, thus enabling the terminal device to obtain the graffiti image. This application does not limit the method by which the terminal device obtains the graffiti image.
[0085] Step 304: Based on the graffiti image, obtain the corresponding set of source images.
[0086] Optionally, the terminal device can input the graffiti image into an image generation model to generate a first set of filtered images. Based on the image parameters of each image in the first set of filtered images, the similarity between each image in the first set and the graffiti image is obtained. Based on the similarity of each image in the first set of filtered images, images with similarity higher than a preset similarity threshold are selected from the first set of filtered images as the source image set. The image generation model can be trained based on a Generative Adversarial Network (GAN). The terminal device inputs the acquired graffiti image into the image generation model, which generates multiple images based on the graffiti image; the set of these images constitutes the first set of filtered images in this step.
[0087] After obtaining the first set of filtered images, the terminal device calculates the similarity between each image and the graffiti image based on the image parameters of each image. Optionally, the image parameters can be parameters such as the outline length, outline shape, and color of the target object contained in each image. For example, if the terminal device obtains three images in the first set of filtered images using the image generation model described above, the terminal device can identify the target object (which can be an object, animal, or person contained in the image) in each image and obtain the outline length, outline shape, and color of the target object. It then calculates the similarity between the two images using the outline length, outline shape, and color of the target object, as well as the outline length, outline shape, and color of the target object in the graffiti image. Similarly, the terminal device can obtain the similarity of each of these three images and obtain the source image set according to the order of these similarity scores.
[0088] Optionally, the terminal device can pre-set a preset similarity threshold. The terminal device uses this preset similarity threshold to filter the similarity of each image in the first set of filtered images, and selects the images with similarity greater than the preset similarity threshold as images in the source image set. For example, if the similarity of the above three images is 90%, 85%, and 70% respectively, and the preset similarity threshold is 75%, then the terminal device can select photos with similarity of 90% and 85% respectively from the above three images as images in the source image set.
[0089] In one possible implementation, the terminal device inputs the graffiti image into an image generation model to generate a first set of filtered images, and then obtains the number of images in the first set of filtered images. When the number of images in the first set of filtered images is not greater than a preset threshold, the graffiti image is input into an image search model to obtain a second set of filtered images. Each image obtained by the image search model contains its own confidence score. Based on the confidence scores of each image in the second set of filtered images, a set of source images is obtained from the second set of filtered images. When the number of images to be filtered is greater than the preset threshold, the step of obtaining the similarity between each image in the first set of filtered images and the graffiti image is performed based on the image parameters of each image in the first set of filtered images.
[0090] That is, the terminal device can also test the results of the image generation model. If the image generation model cannot generate more than a preset threshold number of images, the terminal device can use an image search model to obtain a set of source images. The image search model is based on a modified GitLab open-source project. For example, with a preset threshold of 2, when the number of images in the first set of filtered images generated by the image generation model is less than 2, the terminal device inputs the doodle image into the image search model to obtain a second set of filtered images. Optionally, the terminal device can pre-set a second threshold. The terminal device uses this second threshold to filter the confidence scores of each image in the second set of filtered images, and obtains images with confidence scores greater than the second threshold as images in the source image set. For example, if the confidence scores of the three images mentioned above are 92%, 89%, and 76%, and the second threshold is 80%, then the terminal device can select photos with similarity scores of 92% and 89% as images in the source image set. When the number of each image in the first set of filtered images generated by the image generation model is greater than 2, the terminal device executes the step of obtaining the similarity between each image in the first set of filtered images and the graffiti image based on the image parameters of each image in the first set of filtered images.
[0091] Step 305: In response to the selection operation of the first image, obtain the category information of the first image. The first image is any one of the material images in the set, and the category information is used to indicate the category of the target object in the first image.
[0092] Optionally, after the terminal device obtains the set of source images, the user can independently select any image from it. The terminal device responds to the user's touch operation on the selection control in the functional area, acquires the first image, and then obtains the category information of the first image. For example, in the above... Figure 4In the process, the terminal device can display each image from the acquired image set in the functional area 403. The functional area 403 also includes a selection control 403c. The user can trigger the selection control 403c to select an image from the image set, and the terminal device can obtain the corresponding category information for the selected image.
[0093] Optionally, the category information can be category information such as cat, horse, dog, etc., or category information such as table, chair, computer, etc. That is, the terminal device can obtain the category of the target object in the selected image. Optionally, if the first image is obtained through an image generation model, the category information can be obtained by the terminal device from the image generation model; if the first image is obtained through an image search model, it can also be obtained from the image search model. Alternatively, the terminal device can also input the first image into a separate type recognition model, and use the type recognition model to identify the first image and obtain its category information.
[0094] Step 306: Based on the category information, push the target content of the first image. The target content is used to describe the target object of the first image.
[0095] After obtaining the category information, the terminal device obtains the target content of the first image based on the category information. The target content is the data content corresponding to the target object in the first image.
[0096] Optionally, the drawing interface also includes a push area. The terminal device retrieves the target content of the first image based on category information and pushes the target content into the push area according to a preset push method. For example, in the above... Figure 4 In the drawing interface 400, a push area 404 is also included. This push area 404 contains first target content 404a, second target content 404b, and third target content 404c. The target content of the first image obtained by the terminal device through category information can be displayed in the push area according to a preset push method, thereby achieving push functionality. The preset push method can be based on the order in which the target content is obtained, or it can be based on the update time of the obtained target content, or it can be based on the number of characters in the obtained target content.
[0097] In one possible implementation, the terminal device can also respond to a trigger operation on the target content by displaying a content explanation interface; and explain the target content of the first image according to a preset explanation method. For example, in the above... Figure 4Within the push area (404 error zone), each target content item is clickable. Clicking on any target content item will display a content explanation interface, where the terminal device will explain the target content according to a preset explanation method. Optionally, the preset explanation method can be video, audio playback, text display, etc. Please refer to [reference needed]. Figure 5 This illustration shows a schematic diagram of a content explanation interface according to an exemplary embodiment of this application. Figure 5 As shown, the content explanation interface 500 includes the target content 501, which can be played by the terminal device via voice playback, making it convenient for users to understand the target content.
[0098] In one possible implementation, as described above Figure 4 The functional area 403 also includes an explanation control 403d. Users can select target content within the push area 404 and trigger the explanation control 403d, which can also cause the terminal device to display the above content on the screen. Figure 5 The interface shown explains the target content.
[0099] In one possible implementation, the user can also draw a doodle image containing more than two target objects in the above-mentioned doodle interface. In this case, after obtaining the doodle image, the terminal device can obtain the corresponding material image set for each target object in the doodle image, thereby executing the subsequent steps 305 to 306.
[0100] For example, please refer to Figure 6 This illustration shows a schematic diagram of a graffiti interface according to an exemplary embodiment of this application. Figure 6 As shown, the graffiti interface 600 includes a drawing area 601, a tool area 602, a function area 603, and a push area 604. The function area 603 further includes a shooting control 603a, a "guess" control 603b, a selection control 603c, an explanation control 603d, and a group recognition control 603e. Optionally, when the user draws target objects including first objects 601a and 601b in the drawing area 601, the user can click the group recognition control 603e in the function area 603. The terminal device can then acquire the graffiti image from the drawing area 601, segment the graffiti image (each segmented graffiti image containing a target object), input the segmented graffiti images into an image generation model, obtain the first set of filtered images for each segmented graffiti image, and execute subsequent steps until the target content of each target object is pushed into the push area 604.
[0101] In summary, this application acquires graffiti images via a terminal device; obtains a set of corresponding source images based on the graffiti images; in response to a selection operation on a first image, it acquires the category information of the first image, where the first image is any one of the source images, and the category information indicates the category of the target object in the first image; and based on the category information, it pushes the target content of the first image, where the target content describes the target object of the first image. This application acquires a set of source images from graffiti images, allows selection of one image through human-computer interaction, acquires the category information of the selected image, and then acquires and pushes the target content of the selected image based on this category information. This makes it easier to understand the graffiti images drawn by the user and improves the diversity and accuracy of the recognition results.
[0102] In addition to obtaining the first set of filtered images through an image generation model and then obtaining the material image set, this application also adds a branch of an image search model. This branch can use the result of the image search model when the image generation model fails to convert, which can improve the efficiency of obtaining the content of graffiti images.
[0103] In addition, this application can also recognize graffiti images of multiple target objects, quickly obtain the content of user graffiti images, improve the diversity of graffiti image recognition methods, increase the diversity of recognition results, and improve the accuracy of graffiti image recognition.
[0104] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0105] Please refer to Figure 7 This illustration shows a structural block diagram of an information push device 700 provided in an exemplary embodiment of this application. The information push device 700 can be applied to a terminal device. The information push device 700 includes:
[0106] The first acquisition module 701 is used to acquire graffiti images;
[0107] The second acquisition module 702 is used to acquire the material image set corresponding to the graffiti image based on the graffiti image;
[0108] The third acquisition module 703 is used to acquire category information of the first image in response to the selection operation of the first image, wherein the first image is any one of the material images in the set, and the category information is used to indicate the category of the target object in the first image;
[0109] The information push module 704 is used to push target content of the first image according to the category information, wherein the target content is used to describe the target object of the first image.
[0110] In summary, this application acquires graffiti images via a terminal device; obtains a set of corresponding source images based on the graffiti images; in response to a selection operation on a first image, it acquires the category information of the first image, where the first image is any one of the source images, and the category information indicates the category of the target object in the first image; and based on the category information, it pushes the target content of the first image, where the target content describes the target object of the first image. This application acquires a set of source images from graffiti images, allows selection of one image through human-computer interaction, acquires the category information of the selected image, and then acquires and pushes the target content of the selected image based on this category information. This makes it easier to understand the graffiti images drawn by the user and improves the diversity and accuracy of the recognition results.
[0111] Optionally, the terminal device includes a display screen displaying a drawing interface, the drawing interface including functional areas; the third acquisition module 703 includes: a first acquisition unit and a second acquisition unit;
[0112] The first acquisition unit is configured to acquire the first image in response to a touch operation on a selection control in the functional area;
[0113] The second acquisition unit is used to acquire category information of the first image based on the first image.
[0114] Optionally, the graffiti interface also includes a push area, and the information push module 704 includes: a third acquisition unit and a push unit;
[0115] The third acquisition unit is used to acquire the target content of the first image based on the category information;
[0116] The push unit is used to push the target content in the push area according to a preset push method.
[0117] Optionally, the terminal device includes a display screen, and the apparatus further includes:
[0118] The first display module is used to display the graffiti interface on the display screen before the graffiti image is acquired. The graffiti interface also includes a drawing area and a tool area.
[0119] The first generation module is used to generate a doodle image in response to a drawing operation in the drawing area;
[0120] The first acquisition module 701 is used to acquire the graffiti image generated in the graffiti interface; or, in response to the triggering operation of the shooting control in the functional area, to acquire the imported graffiti image.
[0121] Optionally, the second acquisition module 702 includes:
[0122] The second generation module is used to input the graffiti image into the image generation model to generate the first set of filtered images;
[0123] The fourth acquisition module is used to acquire the similarity between each image in the first filtered image set and the graffiti image based on the image parameters of each image in the first filtered image set.
[0124] The fifth acquisition module is used to select images from the first filtered image set whose similarity is higher than a preset similarity threshold as the material image set, based on the similarity of each image in the first filtered image set.
[0125] Optionally, the device further includes:
[0126] The fifth acquisition module is used to acquire the number of images in the first filter image set after the graffiti image is input into the image generation model to generate the first filter image set;
[0127] The sixth acquisition module is used to input the graffiti image into the image search model to obtain the second filter image set when the number of the first filtered image set is not greater than a preset number threshold. Each image obtained by the image search model contains its own confidence level.
[0128] The seventh acquisition module is used to acquire the material image set from the second filtered image set based on the confidence level of each image in the second filtered image set.
[0129] The first execution module is used to execute the step of obtaining the similarity between each image in the first set of filtered images and the graffiti image based on the image parameters of each image in the first set of filtered images when the number of images to be filtered is greater than the preset number threshold.
[0130] Optionally, the device further includes:
[0131] The second display module is used to display a content explanation interface in response to a triggering operation on the target content after the target content of the first image is pushed according to the category information.
[0132] The content explanation module is used to explain the target content of the first image according to a preset explanation method.
[0133] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an exemplary embodiment of this application. For example... Figure 8 As shown, the terminal device 800 includes a Central Processing Unit (CPU) 801, a system memory 804 including Random Access Memory (RAM) 802 and Read Only Memory (ROM) 803, and a system bus 805 connecting the system memory 804 and the CPU 801. The terminal device 800 also includes a Basic Input / Output System (I / O System) 806 to facilitate information transfer between various devices within the computer, and a mass storage device 807 for storing the operating system 812, application programs 813, and other program modules 814.
[0134] The basic input / output system 806 includes a display 806 for displaying information and an input device 809 for user input, such as a mouse or keyboard. Both the display 806 and the input device 809 are connected to the central processing unit 801 via an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include the input / output controller 810 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, printer, or other types of output devices.
[0135] The mass storage device 807 is connected to the central processing unit 801 via a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable media provide non-volatile storage for the terminal device 800. That is, the mass storage device 807 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0136] The computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the above-mentioned types. The system memory 804 and mass storage device 807 described above can be collectively referred to as memory.
[0137] Terminal device 800 can be connected to the Internet or other network devices via network interface unit 811 connected to the system bus 805.
[0138] The memory also includes one or more programs stored in the memory. The central processing unit 801 executes the one or more programs to implement all or part of the steps performed by the terminal device in the methods provided in the above embodiments of this application.
[0139] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0140] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, terminal device, or data center to another website, computer, terminal device, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a terminal device or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., Digital Video Disc (DVD)), or a semiconductor medium (e.g., Solid State Disk (SSD)). This application also discloses a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method described in the above method embodiments.
[0141] This application also discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the methods described in the above method embodiments.
[0142] This application also discloses an application publishing platform, which is used to publish computer program products. When the computer program products are run on a computer, the computer executes the methods described in the above method embodiments.
[0143] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0144] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-accessible memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests to cause a terminal device (which can be a personal computer, terminal device, or network device, specifically a processor in the terminal device) to execute some or all of the steps of the methods described in the various embodiments of this application.
[0145] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0146] The above provides examples of an information push method, apparatus, terminal device, and storage medium disclosed in the embodiments of this application. These examples illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are merely for the purpose of helping to understand the method and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An information push method, characterized in that, The method is applied to a terminal device, and the method includes: Get the graffiti image; The graffiti image is input into the image generation model to generate a first set of filtered images; Based on the image parameters of each image in the first filtered image set, obtain the similarity between each image in the first filtered image set and the graffiti image; Based on the similarity of each image in the first filtered image set, select images from the first filtered image set whose similarity is higher than a preset similarity threshold as the source image set; In response to the selection operation of the first image, the category information of the first image is obtained. The first image is any one of the material images in the set. The category information is used to indicate the category of the target object in the first image. Based on the category information, target content is pushed to the first image, whereby the target content describes the target object of the first image.
2. The method according to claim 1, characterized in that, The terminal device includes a display screen displaying a drawing interface, the drawing interface including functional areas; the step of obtaining category information of the first image in response to a selection operation of the first image includes: In response to a touch operation on a selection control in the functional area, the first image is acquired; Based on the first image, obtain the category information of the first image.
3. The method according to claim 2, characterized in that, The graffiti interface also includes a push area, wherein pushing the target content of the first image according to the category information includes: Based on the category information, obtain the target content of the first image; The target content is pushed to the push area according to a preset push method.
4. The method according to claim 2, characterized in that, Before obtaining the graffiti image, the following is also included: The drawing interface is displayed on the screen, and the drawing interface also includes a drawing area and a tool area; A doodle image is generated in response to a drawing operation in the drawing area; The process of obtaining the graffiti image includes: Obtain the graffiti image generated in the graffiti interface; or, in response to a trigger operation on the camera control in the functional area, obtain the imported graffiti image.
5. The method according to claim 1, characterized in that, After inputting the graffiti image into the image generation model to generate the first set of filtered images, the method further includes: Get the number of images in the first filtered image set; When the number of images in the first filtered image set is not greater than a preset threshold, the graffiti image is input into the image search model to obtain the second filtered image set. Each image obtained by the image search model contains its own confidence level. Based on the confidence level of each image in the second filtered image set, the material image set is obtained from the second filtered image set; When the number of images in the first filtered image set is greater than the preset number threshold, the step of obtaining the similarity between each image in the first filtered image set and the graffiti image based on the image parameters of each image in the first filtered image set is executed.
6. The method according to any one of claims 1 to 5, characterized in that, After pushing the target content of the first image according to the category information, the method further includes: In response to a trigger operation on the target content, a content explanation interface is displayed; The target content of the first image will be explained according to the preset explanation method.
7. An information push device, characterized in that, The device is used in a terminal device, and the device includes: The first acquisition module is used to acquire graffiti images; The second acquisition module is used to input the graffiti image into the image generation model to generate a first set of filtered images; based on the image parameters of each image in the first set of filtered images, to obtain the similarity between each image in the first set of filtered images and the graffiti image; and based on the similarity between each image in the first set of filtered images, to select each image in the first set of filtered images with a similarity higher than a preset similarity threshold as a set of source images. The third acquisition module is used to acquire category information of the first image in response to the selection operation of the first image, wherein the first image is any one of the material images in the set, and the category information is used to indicate the category of the target object in the first image; The information push module is used to push target content of the first image according to the category information, wherein the target content is used to describe the target object of the first image.
8. A terminal device, characterized in that, The terminal device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor enables the processor to implement the information push method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the information push method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image generation method, device and computer readable storage medium
CN108615253A
Method and device for pushing information
CN111767456A