Method, device and equipment for generating information based on image

By acquiring user touch operations on images, identifying target elements, and generating personalized virtual avatars and introductory information, the problem of uninteresting image-generated information in existing technologies is solved, thereby improving user learning interest and experience.

CN121979437APending Publication Date: 2026-05-05ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing methods for generating information from images lack interest and personalization, and cannot generate personalized teaching content based on user needs.

Method used

By acquiring the user's touch operation on the image displayed on the terminal device, the image area to which the operation is directed is determined, the target element is extracted, and a virtual image and introductory information with the same characteristics as the target element are generated using a generative model and displayed on the terminal device.

Benefits of technology

It enables the generation of personalized virtual avatars and introductory information based on user needs, thereby increasing the fun and motivation of learning and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979437A_ABST
    Figure CN121979437A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, device and equipment for generating information based on an image. The scheme can comprise the following steps: acquiring a touch operation executed by a user for an image displayed by the terminal equipment; the method comprises the steps of determining a region image pointed by a touch operation in an image in response to the touch operation; performing element extraction on the regional image, determining a target element contained in the regional image, and based on the target element, generating a virtual image having the same characteristics as the target element and introduction information for the target element by using a generative model; wherein the virtual image and the introduction information can be displayed on the terminal equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and more particularly to a method for generating information based on images. This specification also relates to an apparatus for generating information based on images, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] As society continues to develop, users can learn knowledge from applications in different ways, such as watching animations, viewing short videos, and viewing text and images. However, most knowledge content is composed of fixed animations or images, which cannot display personalized teaching content according to actual needs and lacks interest.

[0003] Therefore, how to provide a method that can generate personalized information for users to learn or view based on simple interaction is an urgent technical problem to be solved. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide a method, apparatus, device, and computer-readable medium for generating information based on images to address the problem of low interest in existing methods for generating information based on images.

[0005] According to a first aspect of one or more embodiments of this specification, a method for generating information based on an image is provided, comprising: Acquire the touch operations performed by the user in response to the image displayed on the terminal device; In response to the touch operation, determine the region of the image to which the touch operation points in the image; Element extraction is performed on the region image to determine the target elements contained in the region image; Based on the target element, a virtual avatar with the same characteristics as the target element and introductory information about the target element are generated using a generative model; the terminal device is able to display the virtual avatar and the introductory information.

[0006] According to a second aspect of one or more embodiments of this specification, an apparatus based on image generation information is provided, comprising: The operation acquisition module is used to acquire the touch operations performed by the user on the image displayed on the terminal device; A region image determination module is used to determine the region image pointed to by the touch operation in the image in response to the touch operation; An element determination module is used to extract elements from the region image and determine the target elements contained in the region image. An information generation module is used to generate, based on the target element, a virtual image with the same characteristics as the target element and introductory information about the target element using a generative model; the terminal device is capable of displaying the virtual image and the introductory information.

[0007] According to a third aspect of one or more embodiments of this specification, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer instructions, implements the steps of the method for generating information based on an image.

[0008] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions which, when executed by a processor, implement the steps of the method for generating information based on an image.

[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for generating information based on images.

[0010] At least one embodiment of this specification can achieve the following beneficial effects: by acquiring the touch operation performed by the user on the image displayed on the terminal device, determining the area image pointed to by the touch operation in the image, extracting the target element from the area image, and generating a virtual image with the same characteristics as the target element and introductory information about the target element based on the target element using a generative model, the virtual image and introductory information about the target element can be displayed on the terminal. This allows users to generate corresponding introductory information and virtual images for learning and understanding by simply interacting with the image based on their learning needs; it eliminates the need for users to manually search for knowledge related to the target element, and also improves the interest and personalization of the generated information, increases users' enthusiasm for learning and understanding the target element, and improves the user experience. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the overall architecture of a method for generating information based on an image, provided in one embodiment of this specification. Figure 2This is a flowchart illustrating a method for generating information based on an image, as provided in one embodiment of this specification. Figure 3 This is a swimlane diagram of a method for generating information based on an image, as provided in one embodiment of this specification. Figure 4 This is a schematic diagram of an image displayed in a terminal device according to an embodiment of this specification; Figure 5 This is a schematic diagram illustrating a user's touch operation on an image, provided in one embodiment of this specification. Figure 6 This is a schematic diagram showing a virtual image and introductory information provided in one embodiment of this specification; Figure 7 This specification provides an embodiment corresponding to... Figure 2 A schematic diagram of a device based on image generation information; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0013] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0014] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “an,” “an,” “the,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification includes any or all possible combinations of one or more associated listed items.

[0016] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded.

[0017] Although the terms "first," "second," etc., may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of this specification. Ordinal numbers such as "first," "second," etc., do not necessarily indicate order; often they are used to facilitate the distinction of objects. For example, "first server" and "second server" usually refer to two servers. To distinguish these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0018] Depending on the context, the word "if" as used here can be interpreted as "when," "when," or "in response to determination."

[0019] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0020] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0021] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.

[0022] The following explains the terms and concepts used in one or more embodiments of this specification.

[0023] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0024] Figure 1 This is a schematic diagram illustrating the overall architecture of a method for generating information based on images, provided as an embodiment of this specification. Figure 1 As shown, the solution may include a terminal device 1 and a server 2. Terminal device 1 may have a display component for displaying images; terminal device 1 may also sense user operations on the displayed images. Users can perform operations on the images in terminal device 1, and terminal device 1 can send information about the user's operations to server 2. Server 2 can determine the region of the image to which the operation is directed. Server 2 can extract elements from the region of the image to obtain the target element contained in the image. Based on the target element, server 2 can use a generative model to generate a virtual image with the same characteristics as the target element, along with introductory information about the target element. Server 2 sends the generated virtual image and introductory information to terminal device 1, and terminal device 1 can display the virtual image and introductory information. Terminal device 1 can interact with the user through a graphical user interface to invoke the server, thereby implementing the method provided in the embodiments of this specification.

[0025] In such Figure 1In the application scenario shown, server 2 can connect to one or more terminal devices 1 via a local area network (LAN), a wide area network (WAN), an internet connection, or other types of data networks. Figure 1 Server 2 in the context can include, but is not limited to, any device, equipment, platform, or equipment cluster with computing and processing capabilities. Figure 1 The terminal device 1 may include, but is not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices.

[0026] If the terminal device's operating resources can meet the execution conditions for processing the image-based music generation method, it can also be executed by the terminal device.

[0027] This application provides a method for generating information based on an image, and also relates to an apparatus for generating information based on an image, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0028] Figure 2 This is a flowchart illustrating a method for generating information based on an image, as provided in one embodiment of this specification.

[0029] From a programming perspective, the executor of the process can be a program hosted on an application server or application terminal. From a hardware perspective, the executor of the process can be a server or terminal. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.

[0030] like Figure 2 As shown, the process may include the following steps: Step 202: Obtain the touch operation performed by the user on the image displayed on the terminal device.

[0031] In one embodiment of this specification, the image may be acquired by the user using the camera device of the terminal device. Specifically, it may be an image captured by the user using the terminal device, such as by taking a picture or video; or an image scanned by the camera after it is activated, without the user clicking a shooting control, such as a photo button or video recording button. The image may also be selected by the user from the local storage of the terminal device; or, the image may be selected by the user from multiple images provided by the default of the currently displayed application on the terminal device; or, the image may be randomly displayed by the currently displayed application based on multiple default images; or, the image may be generated based on the user's operation on the image control, and so on.

[0032] In practical applications, users can open application pages, H5 pages, or mini-program pages on their terminal devices that generate information based on images. The page will display a shooting interface, which may contain shooting controls. Users can click on these controls to take a picture and obtain an image. The shooting interface includes a sidebar containing controls for shooting, local images, system images, and text-to-image editing. Users can interact with these controls and choose the appropriate method to display the image on the terminal device for further image manipulation. The sidebar can be located anywhere on the left, right, bottom, or top of the interface.

[0033] In practical applications, users can open application pages, H5 pages, or mini-program pages on their terminal devices that generate information based on images. The page will display a default system image. If the user doesn't like the default image, they can use the image selection controls on the page to access a menu bar. This menu bar can include controls for local images, system images, text-based images, scanning, and capturing images, allowing the user to select the appropriate method to change the default image. The default system image displayed each time the page is opened can be different, or it can be the same each time.

[0034] Touch operations can be one or more taps on a specific location in an image; or a long press on a specific location in an image; or a circle drawn on a specific area of ​​an image; or a voice selection of an element within an image, such as inputting a voice message like "introduce the flower" on an image containing flowers, and the server or terminal can then use the flower as the target element.

[0035] Step 204: In response to the touch operation, determine the region image in the image to which the touch operation is directed.

[0036] In one embodiment of this specification, the region image may be a portion of the image containing the location of the touch operation. If a user operates on a certain location in the image, the server can perform image extraction based on the location of the user's operation. For example, the region image can be obtained by expanding outwards by a preset distance from the location of the user's operation in the image, or by combining or determining the region image based on pixel changes. If the user operates on a certain area in the image, that area can be extracted to obtain the region image; or, the server can determine the area of ​​the target element in the image based on the area operated by the user and the pixel information in the adjacent areas of that area, and extract the area of ​​the target element in the image to obtain the region image.

[0037] Step 206: Extract elements from the region image to determine the target elements contained in the region image.

[0038] In one embodiment of this specification, an element can be object information contained in an image, such as trees, rivers, fountains, seats, glasses, animals, etc., contained in the image. A target element can represent the element corresponding to the position of the user's operation in the image.

[0039] If the operation targets a specific location in an image, a predetermined distance can be extended outward from the operation location to obtain an image region. The target element can then be identified based on the image features within this region. Alternatively, if the operation targets a specific area within an image, the target element can be identified within that region. Or, if the operation is a speech operation targeting an element in an image, the image features matching the speech features can be determined based on the speech characteristics of the operation, and the matching image feature region can be identified as the target element.

[0040] In one embodiment of this specification, preprocessing can be performed before element extraction from the region image. Specifically, noise reduction, contrast enhancement, and size standardization can be applied to the region image, such as removing blur and noise, and uniformly adjusting it to 512*512 pixels. This allows for element extraction from the preprocessed region image to obtain the target elements.

[0041] Step 208: Based on the target element, use a generative model to generate a virtual image with the same characteristics as the target element, as well as introductory information about the target element.

[0042] The terminal device can display the virtual avatar and the introductory information.

[0043] In one embodiment of this specification, a virtual avatar can represent a visualized virtual digital image generated based on a target element. This avatar can be a representation of the target element (e.g., a representational / anthropomorphic / symbolic representation), or an IP character from a movie, TV series, or game. For example, if the target element is a "rabbit," a rabbit wearing children's clothing and waving can be generated; or a highly realistic rabbit image identical to a real rabbit can be generated; or a line drawing of a rabbit can be generated, and so on. The virtual avatar can be a 2D or 3D image. The virtual avatar can have actions; these actions can be randomly generated or specified. The virtual avatar can have positive expressions; these expressions can be randomly generated or specified. The virtual avatar can also be an image generated based on a user-specified type, such as a user-specified anthropomorphic image.

[0044] In one embodiment of this specification, the introductory information may represent structured information formed by organizing the core attributes, functions, background, usage methods, etc. of the target element, and is a textual / audio interpretation of the target element. For example, if the target element is a sunflower, the introductory information may include: "Name: Sunflower, Family and Genus: Asteraceae, Helianthus, Flowering Period: July-September, Growth Habit: Sun-loving".

[0045] In one embodiment of this specification, the generative model can be an existing model capable of generating virtual avatars based on images, such as ChatGPT-4, ElevenLabs, Runway Gen-2, etc.; or it can be trained based on requirements. The generative model can extract element features of target elements from a region image and generate corresponding virtual avatars based on these element features, so that the virtual avatars have the same or similar features as the target elements, such as the same shape features, the same color features, etc.

[0046] In one embodiment of this specification, the server can extract visual features of target elements from an image, such as at least one visual attribute like shape, color, texture, and proportion. Based on these visual attributes, the server can determine the basic attributes of the target element, and based on these basic attributes, determine relevant information. The basic attributes and relevant information are then used as introductory information. A virtual avatar is generated based on the basic attributes and relevant information. The terminal device can render the virtual avatar and introductory information according to a preset template and display them on the terminal device's interface.

[0047] In one or more embodiments of this specification, the server can generate a virtual avatar that conforms to the style represented by the style tag selected by the user, thereby increasing user interest and experience. Optionally, generating a virtual avatar with the same characteristics as the target element may include: obtaining a style tag; the style tag being used to instruct the generative model to generate a virtual avatar that conforms to the style corresponding to the style tag; the style tag including at least one of cartoon style, 2D style, 3D style, line drawing style, ink painting style, anthropomorphic style, and fairy tale style; providing the target element and the style tag to the generative model, and using the generative model to obtain a virtual avatar with the same characteristics as the target element and conforming to the style corresponding to the style tag.

[0048] In one embodiment of this specification, the interface of the terminal device displaying the image may include multiple different style tags that can be selected by the user. If the user selects a style tag, the terminal device can sense the selected style tag and send it to the server, enabling the server to generate a virtual avatar that matches the style tag after receiving it. In addition to the style tags provided by the terminal device, the user can also input custom style tags. If the user inputs a custom style tag, they can briefly describe its characteristics so that the server can generate a virtual avatar with corresponding characteristics based on those characteristics. In practical applications, the server can also count the number of times the same custom style tag is input. When the number of inputs reaches a preset value, such as 10 times or 100 times, the custom style tag can be added to the database, allowing the terminal device to display the newly added style tag on the image display interface.

[0049] Generative models have the ability to generate virtual avatars with different styles based on different style tags. Specifically, generative models can determine the style characteristics of style tags based on style tags and generate virtual avatars that conform to the style characteristics.

[0050] In practical applications, generative models can be trained on large multimodal models; alternatively, they can be trained on neural network models. Specifically, sample image information for different elements of different styles can be collected in advance; virtual avatars corresponding to each sample image information can be obtained, constructing sample pairs of virtual avatar-sample image information-style labels; the model can be trained using these sample pairs to obtain the trained generative model. The virtual avatars can be hand-drawn images based on image information by an artist, or images generated based on text information using a text-to-image model.

[0051] In practical applications, the server can also generate virtual avatars based on preset rules. The server can pre-build a mapping library of element features and style tags to generation rules, obtain the element features and style tags of the target element, retrieve the corresponding generation rules from the mapping library, and generate the virtual avatar based on these rules. Specifically, the avatar style can be determined based on the element category and style tags. For example, an entity object of "rabbit" + cartoon style can be mapped to a cartoon chibi style; similarly, an entity object of "geometric shape" + figurative style can be mapped to a figurative geometric style. The core form of the avatar is determined based on visual features, such as determining the color and ear shape of the generated cartoon avatar based on the rabbit's fur color, eye color, and ear shape. After determining the corresponding avatar style and core form based on element features, the server can call image synthesis algorithms or drawing tools to construct the virtual avatar.

[0052] In one or more embodiments of this specification, element features of a target element can be identified, and a corresponding virtual avatar can be generated based on these element features to improve the fit between the virtual avatar and the target element. Optionally, the method may further include: identifying at least one feature information among the name, category, and material of the target element; providing the target element and the style tag to the generative model includes: providing the feature information and the style tag to the generative model, wherein the generative model can generate a virtual avatar with the same features as the target element and conforming to the style corresponding to the style tag based on the feature information and the style tag.

[0053] In one embodiment of this specification, element features can be extracted from the image using a multimodal model, such as using at least one of the models Material Palette, Qwen 2.5 VL, MobileNet, and ResNet to extract features from elements contained in the image. Material can represent the material category of the element; for example, a seat can include wood, iron, alloy, plastic, etc., or material can represent physical properties such as the object's hardness, density, elasticity, and structural form. Element category can represent the category to which the element belongs, such as teaching tools, scenery, etc.; or a more granular category, such as tree, grass, fruit, person, car, etc. Element name can represent the name information of the actual element, such as white rabbit, calico cat, pine cone, apple, etc.

[0054] In one embodiment of this specification, the server can identify the element features of the target element using preset rules, such as preset correspondences between the visual features of each element and its material, thereby determining the corresponding material based on the visual features obtained from the image. Alternatively, the server can also use a neural network model or a large model to identify the element features of the target element. Specifically, a model with the function of identifying materials, categories, or names from images can be used to identify the target element and obtain its element features. This can be set according to actual needs and is not intended as a specific limitation.

[0055] In one embodiment of this specification, a generative model can generate a virtual image with element feature information and conforming to style tags. For example, if the element is a bench, the style tag is 3D style, and the feature information can include the feature information of wood, then a puppet-style 3D virtual image can be generated.

[0056] In practical applications, the server can also acquire scene information, inputting the scene information, style tags, and feature information into the generative model to obtain the virtual avatar output by the generative model. Scene information can be determined by the user's selection of scene tags displayed on the terminal device's interface; alternatively, it can be obtained by the server through image recognition; or it can be the default scene of the application on the terminal device, such as a children's scene. This allows for the generation of virtual avatars that better match the style and features represented by the scene and style tags, based on the scene information and feature information.

[0057] In practical applications, virtual avatars can be generated based on scene information and the feature information of target elements; or, virtual avatars can be generated based on style tags, the element features of target elements and scene information; or, virtual avatars can be generated based on style tags and the element features of target elements.

[0058] In one or more embodiments of this specification, the server may further generate virtual avatars by combining user characteristics, making the generated virtual avatars more in line with user preferences and improving user experience. Optionally, the method may further include: obtaining user characteristic information of the terminal device; the user characteristic information includes the user's historical interaction preference information; the historical interaction preference information represents the interaction information generated by the user interacting with an IP; determining the IP avatar preferred by the user based on the user characteristic information; generating a virtual avatar with the same characteristics as the target element includes: providing information about the target element and the IP avatar preferred by the user to a generative model, and using the generative model to obtain a virtual avatar with the same characteristics as the target element and conforming to the image style of the IP avatar; the information about the IP avatar preferred by the user includes at least one of text information described in natural language and an image of the IP avatar; the text information includes at least one of the name of the IP avatar, visual information of the IP avatar, and category information of the IP avatar.

[0059] In one embodiment of this specification, user characteristic information can be feature data reflecting attributes such as user identity, behavior, and interests. User characteristics can include static features, dynamic behavioral features, and scene-related features. Static features can include features such as occupation and geographical location. Dynamic behavioral features can include historical painting viewing records, such as cartoon style, 2D style, 3D style, or line art style; dynamic behavioral features can also include interactive feedback features, such as collecting, sharing, or liking a certain type of painting; dynamic behavioral features can also include the user's purchase records, such as purchase records for a certain type of item; dynamic behavioral features can also include the user's browsing records, such as viewing records for a certain video.

[0060] In one embodiment of this specification, the server can obtain historical interaction preference information of the user regarding IP addresses from channels such as user registration information, preference questionnaire information, automatically recorded user behavior, or third-party platforms authorized by the user. Specifically, the server can obtain user interaction information for various IP addresses; for any given IP address, count the number of interaction messages for that IP address; and determine the interaction information corresponding to the IP address with the largest number of interactions as historical interaction preference information. Alternatively, the server can obtain user interaction information for various IP addresses; determine the weight of each interaction type in the interaction information; for any given IP address, perform a weighted sum based on the weights and quantities of various interaction information to determine the user bias for that IP address; and determine the interaction information corresponding to the IP address with the highest bias as historical interaction preference information.

[0061] In one embodiment of this specification, an IP character can be a character or symbol that is creatively designed to possess a unique personality, narrative connotation, and emotional value, capable of evoking audience resonance and being commercially developed. IP characters can include many types, such as IP characters from animation and games, IP characters from film and television dramas, IP characters from brand mascots, original trendy IP characters, and IP characters based on real people.

[0062] In practical applications, servers can use user profile information to determine which IP-related products a user has purchased, thus identifying the IP's image. Alternatively, servers can use user profile information to identify characters in anime or movies that a user enjoys watching, and then identify the corresponding IP image based on those characters.

[0063] In practical applications, generative models can acquire the element features of the target element and the image features of the IP image. Based on the element features and image features, virtual images that combine the features of both can be generated, so that the virtual images have both the characteristics of the IP image preferred by users and the characteristics of the target element, thereby improving the user experience.

[0064] In one embodiment of this specification, the server can pre-determine structured information for the IP character, and this structured information is input into the generative model as information about the IP character. Specifically, the server can determine the visual style information of the IP character, including its art style, such as flat cartoon, 3D realistic, ink painting style, or cute chibi style; the visual style information can also include color schemes, such as black and white, or red, white and blue; the visual style information can also include line features, such as thick lines, fine lines, or no outline; the visual style features can also include detailed elements, such as IP-specific accessories. The server can also determine the morphological specification information of the IP character, including its proportional features, such as a head-to-body ratio of 1:1 or 7:1; the morphological specification information can also include styling logic, such as a round head with triangular ears and a cylindrical body shape. The server can determine the personality association information of the IP character so as to generate virtual characters with expressions or actions associated with the personality. The personality association information can include the personality traits of the IP character, such as innocent and cute, cool and handsome, or gentle and healing; the personality association information can also include classic expressions or classic action images of the IP character. Structured information can include at least one of the aforementioned visual style information, morphological specification information, and personality-related information. The server can generate virtual characters with the same or similar characteristics as the IP character, fitting the IP's style and personality. For example, if the target element is a poplar tree, the IP character's structured information includes visual style (cute chibi style, blue and white color scheme, rounded lines, no rough texture), morphological specification (1:1 head-to-body ratio, rounded body), and exclusive elements (crown accessories, innocent expression). After the model identifies the core structure of the poplar tree—"straight trunk + upward-facing fan-shaped crown"—it can generate a virtual poplar tree character with the poplar structure as its framework, a trunk drawn with rounded lines, a crown filled with a blue-green gradient, a mini crown added to the top of the trunk, and overall proportions adjusted to a chibi IP style. This results in a virtual poplar tree character that possesses both the characteristics of a poplar tree and the characteristics of the IP character.

[0065] In practical applications, the server can identify one or more popular IPs based on the feature information of multiple users, provide the IP image information and target elements of a popular IP to the generative model, so that the generative model can generate a virtual image that conforms to the image style of the popular IP, and so on.

[0066] In one embodiment of this specification, the text information described in natural language can be generated by the server processing the image of the IP character using a graph-based text model. Specifically, the server can determine the image corresponding to the IP character and generate the corresponding text information using the graph-based text model, or it can be extracted from existing text descriptions. The graph-based text model can be at least one of ERNIE-ViLG, Qwen-VL, BLIP, GIT, CLIP, etc. The text information can also be natural language description information entered by the user on the interface of the terminal device for a preferred IP character. Specifically, the user can operate on the image requirement control displayed on the interface of the terminal device, and the terminal device can respond to the user's operation by displaying an input box, in which the user can enter a natural language description of the IP character. In practical applications, the user can also enter an image of the IP character in the input box. To avoid the user setting the IP character every time a virtual character is generated for a target element, the text information or the image of the IP character entered by the user can be used as a fixed reference image. If the user wants to use another IP character as a reference for generating a virtual character, they can operate on the image requirement control again to change the IP character to a new image.

[0067] In practical applications, if the IP image is determined by the server based on user characteristics, the server can generate virtual images for different elements using the same IP image within a preset time period. After the preset time period, the server can re-collect user characteristics to determine a new IP image for generating virtual images for different elements. The new IP image can be the same as the old IP image. Alternatively, the server can obtain the IP image based on user characteristics each time a virtual image is generated to generate a virtual image for the target element.

[0068] In one embodiment of this specification, the image of the IP character can be obtained by the server from a video or image based on the character name involved in the product purchased by the user. For example, if the user purchases a 75mm badge of character A, the character image displayed in the product can be used as the corresponding IP character image. Alternatively, the server can determine that character A is a character in the game W, and can search for related images and videos on web pages or applications using keywords to obtain the IP character image of character A. The server can also determine the IP character image based on the video or image viewed by the user. If it is a video, an image containing the IP character can be extracted from the video.

[0069] In practical applications, the terminal device interface can include an IP control. Users can select or deselect the IP control. If the user selects the IP control, it indicates that a virtual avatar based on the IP image should be generated; if the user deselects the IP control, it indicates that a virtual avatar should not be generated. This allows users to freely choose whether or not to generate a virtual avatar based on their IP image, improving the user experience.

[0070] As one implementation method, the server can provide the IP character image and target elements to the generative model, and instruct the generative model what content to generate. The instructions provided by the server to the generative model can be in at least one of the following forms: text or visual annotation. The instructions can be user input obtained by the server to constrain the generation target; or they can be instructions generated by the server based on style tags pre-selected by the user. The instructions can be constraints used to clarify the generation target, specifically including at least one textual form of instructions such as style type (e.g., cartoon, minimalist, realistic, specific art style), form type (e.g., anthropomorphic modification, head-to-body ratio, detail modification), and the generation dimension of the virtual character (e.g., 2D, 3D, static, dynamic). Visual annotations can be constraints for generating the virtual character annotated in the IP character image or target elements. For example, selecting the eyes of the IP character and annotating "use the IP character's eyes as the virtual character's eyes"; selecting the clothing of the IP character and annotating "use the IP character's clothing as the virtual character's clothing"; selecting a certain area of ​​the target element and annotating "use the color of this element area as the main color of the virtual character," etc. This enables generative models to generate virtual avatars that match the IP image and target elements, based on the indicated content and have user-customizable features.

[0071] As another implementation, the server can provide the generative model with text information describing the IP image and target elements, and instruct the generative model what content to generate. The instructions provided by the server to the generative model can be in the form of text or visual annotations, at least one of these. The generative model can generate personalized virtual images based on text and image information.

[0072] In one embodiment of this specification, if the target element is input into the generative model, the generative model can extract the element features for the target element. The generative model can also perform encoding processing on the text information containing the indication content to obtain semantic features; after aligning the semantic features and the element features, a virtual image is generated based on the aligned semantic features and element features. In practical applications, the server can also convert the target element into text information, input the text information used to describe the target element into the generative model, and obtain a virtual image. Specifically, whether the input to the generative model is in the form of pure text, graphic text, or picture is not limited here and can be set based on actual needs.

[0073] In practical applications, the information of the IP image preferred by the user and the regional image can also be input into the generative model; the generative model can extract image features from the regional image, and the image features can specifically include at least one of low-order features (such as contours, edges, color distributions, etc.), middle-order features (such as object structures or human postures, etc.), and high-order features (such as style information, semantic information, etc.). The generative model can determine the IP image features based on the information of the IP image preferred by the user; a virtual image is generated based on the IP image features and the image features.

[0074] In one embodiment of this specification, the text information can be the introduction information presented in the form of natural language description. Specifically, it can be a smooth paragraph description, such as when the user performs a click operation on the area showing a tree in an image, assuming the tree is a poplar tree, the terminal can display text introduction information about the tree such as "The poplar tree is a tall deciduous tree with a straight trunk, widely distributed in temperate regions, and is a common tree species for greening"; or, it can also be structured text, such as "[Name] - Poplar Tree, [Family and Genus] - Compositae, [Flowering Period] - July to September"; or, it can also be concise text, such as "Chinese: Poplar Tree, English: Poplar".

[0075] In one embodiment of this specification, the audio information can be the introduction information presented in the form of voice broadcast with sound as the carrier. Specifically, it can be the form of pure voice broadcast to read the text content; or, it can also be audio with background sound, such as the form of background music adapted to the scene + voice broadcast for the text content; or, it can also be dialogical audio, such as "A: What kind of tree is this? B: "This is a poplar tree, with a very straight trunk and has advantages such as fast growth rate and strong drought tolerance!" In one embodiment of this specification, the video information can be the introduction information presented intuitively and three-dimensionally with dynamic images as the carrier, combined with information such as pictures, sounds, and texts. Specifically, the video information can be at least one type of video such as an explanatory video, an animated video, or a live-action + subtitle video.

[0076] As one embodiment, the server can pre-establish a mapping relationship library containing various text description information and various elements, or a mapping relationship library containing various text description templates and various elements; the server can determine the corresponding text description information based on the target element for display; or the server can determine the corresponding text description template based on the element type of the target element, and fill the text description template for the target element to obtain the text description information.

[0077] As another embodiment, the server can also pre-establish a mapping relationship library containing various audio introduction information and various elements, or a mapping relationship library containing various audio introduction templates and various elements; so that the server can generate corresponding audio introduction information based on the target element and the mapping relationship library.

[0078] As another embodiment, the server can also pre-establish a mapping relationship library containing various video introduction information and various elements, or a mapping relationship library containing various video introduction templates and various elements; so that the server can generate corresponding video introduction information based on the target element and the mapping relationship library.

[0079] As another embodiment, the server can also utilize a multimodal large model to generate descriptive information for the target element, such as Gemini Ultra, GPT-4o, and Canva models. Based on the user-selected descriptive type tag, the server can prompt the multimodal large model to generate the corresponding type of descriptive information. For example, if the user selects the "text" tag, the server can input a prompt indicating the generation of text descriptive information for the target element after inputting the image or text description information of the target element into the multimodal large model. This allows the multimodal large model to generate the target text descriptive information based on the target element's relevant information and the prompt content.

[0080] In one or more embodiments of this specification, the introductory information may include name information of the target element presented in multiple language types, or descriptive information regarding the usage and / or growth of the target element.

[0081] In one embodiment of this specification, the language type can be languages ​​from different countries or regions, such as Chinese, English, Russian, French, German, and Spanish. The language type can also include dialect types and standard language types of various countries. Usage descriptions can be natural language descriptions of the purpose, usage, and applicable scenarios of a target element with practical functions. For example, poplar trees can be used to make furniture, paper, and building materials. Growth descriptions can be natural language descriptions of the growth cycle, growth conditions, and morphological changes of a target element with life characteristics. For example, poplar trees: seedling stage 1-2 years, maturity stage 5-8 years, lifespan up to 20-30 years; growth conditions: suitable for temperate regions; morphological changes: seedling stage: slender trunk, small and sparse leaves; mature stage: thick and straight trunk, fan-shaped crown, willow-leaf-shaped leaves; growth characteristics: rapid growth rate. This allows users to learn about the target element by generating introductory information.

[0082] In one or more embodiments of this specification, introductory information that matches the user's cognition can also be generated, enabling the user to understand the introductory information and improving the user experience. Optionally, the method may further include: determining the user's cognitive level based on the user characteristic information; and generating introductory information for the target element that matches the cognitive level.

[0083] In one embodiment of this specification, the cognitive level can be the depth of knowledge expression defined according to child development psychology (such as Piaget's theory of cognitive development stages). The cognitive level can be adapted to the user's educational stage or age. Taking educational stage as an example, for toddlers, the focus is on sensory descriptions ("Pine cones feel hard, like little hedgehogs"), while for school-aged children, causal logic is introduced ("Pine cones protect the seeds; when they dry, the scales open, and the seeds fly away"). Taking age as an example, if the user is a 3-5 year old child, the information presented focuses on sensory descriptions such as shape, color, and sound; if the user is 6-8 years old, logical content such as ecological relationships and uses is added.

[0084] In practical applications, cognitive levels can be extracted from a knowledge graph and matched to the user's educational stage or age. A knowledge graph can be a structured semantic network. It can be a structured network used to represent entities (such as "pine cone") and their attributes (such as "belongs to pine cones," "edible," "growth cycle") and relationships (such as "eaten by squirrels," "ripens in autumn"). The knowledge graph in this invention can be constructed hierarchically according to the user's age or educational stage (such as 3–5 years old, 6–8 years old), with each layer containing knowledge nodes adapted to that cognitive level.

[0085] In one embodiment of this specification, the server can pre-construct a multi-level educational knowledge graph, with each entity node labeled with its applicable age range. Based on the user's registered age or historical interaction behavior, the server can predict the user's cognitive stage. Once a target element (such as a "pine cone") is identified, a subgraph matching the user's age group can be determined from the knowledge graph. The knowledge points extracted from the subgraph are converted into natural language descriptions and stylized prompts (such as "narrated in a fairytale tone") are injected for the generative model to generate introductory text. This allows for the integration of developmental psychology theory with AI-generated content, enabling age-appropriate instruction. This avoids information overload or overly childish content, improves the effectiveness and receptiveness of educational content, supports personalized learning paths, and provides core technological support for intelligent educational terminals.

[0086] In one or more embodiments of this specification, the server may also generate audio information that conforms to the user's cognitive level to improve the user experience. Optionally, the method may further include: if the cognitive level is lower than a preset level, generating audio of the introductory information so that the user terminal plays the audio while displaying the introductory information; or, if the cognitive level is lower than a preset level, generating an introductory video based on the introductory information so that the user terminal plays the introductory information in video format.

[0087] In one embodiment of this specification, the preset level can be set based on the user's age or educational stage; or, the preset level can be determined based on expert experience; or it can be determined according to the type of information being introduced, such as text-based information, where the preset level can be set to the cognitive level corresponding to the third grade of elementary school; when the user's cognitive level is lower than that corresponding to the third grade of elementary school, it can be determined that the user's cognitive level is low and may not be able to understand or recognize the text displayed on the terminal, requiring audio or video assistance to help the user understand the information being introduced.

[0088] In one embodiment of this specification, the audio can be the same as the text content displayed on the terminal device, or the audio can be different from the text content displayed on the terminal device. The video information can be in the form of animation. Therefore, by using audio or video, even users with lower cognitive abilities can understand and learn about the introductory information of the target element, thereby increasing user learning motivation.

[0089] In one or more embodiments of this specification, the auditory style of the audio has the same emotional dimension as the visual style of the virtual character; or, the auditory and visual style of the introductory video has the same emotional dimension as the visual style of the virtual character; the emotional dimension includes at least one of lively, calm, childlike, realistic, humorous, and serious.

[0090] In one embodiment of this specification, the emotional dimension can be used to quantify the emotional tendency conveyed by the virtual avatar. This can typically be mapped to a continuous space using a pre-trained emotion classifier or manual annotation. For example, liveliness and composure are two different emotions; childlike innocence and seriousness are two different emotions. Auditory style can be mapped to visual style. The corresponding auditory style is determined based on the visual style of the virtual avatar, enabling the generation of appropriate videos or audio based on the auditory style. This ensures that the sound in the generated audio or video conforms to the virtual avatar while also possessing the corresponding emotional tendency.

[0091] In one or more embodiments of this specification, the server may also generate a sound that matches the target element, thereby improving the user experience. Optionally, the introductory information for the target element may include audio information; the audio information is generated using a text-to-speech engine according to a voice that conforms to the image characteristics of the virtual avatar; the text-to-speech engine can dynamically adjust the timbre, speech rate, and intonation parameters of the audio information according to the image characteristics of the virtual avatar; the image characteristics include the virtual avatar's life stage characteristics, personality characteristics, and physiological attribute characteristics.

[0092] In one embodiment of this specification, the Text-to-Speech Engine (TTS Engine) can be a core software component capable of converting written text into natural speech output, serving as a concrete implementation carrier for speech synthesis technology. The TTS engine can adjust the timbre, speech rate, and intonation parameters of the audio information based on the virtual avatar's characteristics. For example, if the virtual avatar is lively, the timbre can be a child's timbre, the speech rate can be relatively fast, and the intonation can be relatively upbeat and lively; if the virtual avatar is calm and composed, the timbre can be a middle-aged timbre, the speech rate can be steady, and the intonation can be relatively deep and steady.

[0093] In one embodiment of this specification, virtual avatars of different styles have different timbres. For example, the timbre of a virtual avatar with a child style can be a clear, childlike voice; the timbre of a virtual avatar with a mature style can be a deep, resonant voice; and the timbre of a virtual avatar with a futuristic, technological style can be electronic music.

[0094] Regarding the determination of timbre, specifically, a pre-established timbre library containing the correspondence between image features and timbres can be created. After determining the image features of the target element of the user's operation, timbres that map to the image features can be obtained from the pre-established timbre library. Alternatively, a pre-trained model capable of generating timbres can be used to process the image features, obtaining a timbre output by the model that matches the image features. The model can be a pre-trained neural network model or a pre-trained large model. If it is a large model, the server can generate corresponding prompts based on the extracted image features, and then call the large model to generate a timbre matching the image features based on the prompts, eliminating the need for the user to provide corresponding prompts and improving the user experience. The embodiments in this specification can quickly and accurately obtain timbres matching image features through the above methods, facilitating the subsequent generation of audio information that conforms to the image features of the virtual character.

[0095] In one embodiment of this specification, the sound content in the audio information may be completely consistent with or not completely consistent with the displayed text content. For example, interjections or greetings may be added to the audio; the virtual avatar in the audio information may introduce the target element in the first person.

[0096] In one or more embodiments of this specification, the user can also interact with controls displayed on a page of the terminal device (such as at least one page displaying introductory information, an image, or a virtual avatar), enabling the user to view the introductory information via video or audio, thereby improving the user experience. Optionally, if the introductory information for the target element may include audio information, the page displaying the virtual avatar on the terminal device further includes an audio playback control; if the terminal device receives a trigger operation from the user on the audio playback control, it plays the audio information; or, if the introductory information for the target element includes video information, the page displaying the virtual avatar on the terminal device further includes a video playback control; if the terminal device receives a trigger operation from the user on the video playback control, it plays the video information.

[0097] In one embodiment of this specification, the audio playback control can be operated by the user. The terminal device can respond to the user's operation on the audio playback control and control the audio information based on the state of the audio information. For example, if the audio information is in a playing state, the terminal device can stop playing the audio information after the user operates the audio playback control; if the audio information is in a not playing state, the terminal device can play the audio information after the user operates the audio playback control. The form of the audio playback control can be different depending on the audio information state. The audio playback control needs to be displayed on the page displaying the virtual avatar, and its position can be below the virtual avatar, in the bottom toolbar of the page, or in a floating area next to the virtual avatar, so as not to obscure the core visual aspects of the virtual avatar (such as the face and main shape) while facilitating user viewing and operation.

[0098] In one embodiment of this specification, the user can operate the audio playback control through preset, playback-triggering interactive operations, such as clicking the play button, long-pressing the control, or sliding the progress bar to jump to the playback position. In practical applications, the terminal device's page can also include a volume control control, allowing the user to adjust the volume to avoid the audio being too low and unclear, or too high and negatively impacting the user experience. The terminal device's page can also include a progress bar, which is closely related to the audio playback progress. The user can adjust the progress bar to ensure the audio plays at the appropriate position.

[0099] In one embodiment of this specification, the audio information may be a voice broadcast of introductory information in text form; or it may be a broadcast of English words of the target element; or it may be a broadcast of English sentences composed of English words of the target element; or it may be a broadcast of introductory information in text form and English words of the target element, etc.

[0100] In one embodiment of this specification, if the video information includes the linkage between the screen and the virtual image, such as the virtual image in the video having the same style as the virtual image displayed on the page, the video playback area and the virtual image display area can be merged. For example, the virtual image can appear in the video screen as a narrator to explain the introductory information of the target element, thereby improving visual unity.

[0101] In one embodiment of this specification, the video playback control can be operated by the user. The terminal device can respond to the user's operation on the video playback control and control the audio information based on the audio information's state. For example, if the audio information is in a playing state, the terminal device can stop playing the audio information after the user operates the video playback control; if the audio information is in a non-playing state, the terminal device can play the audio information after the user operates the video playback control. The user can operate the video playback control through preset interactive operations that trigger playback, such as clicking the play button, long-pressing the control, or sliding the progress bar to jump to the playback position. The form of the video playback control can be different depending on the audio information state. The video playback control needs to be displayed on the page displaying the virtual avatar, and its position can be below the virtual avatar, in the bottom toolbar of the page, or in a floating area next to the virtual avatar, so as not to obscure the core visuals of the virtual avatar (such as the face and main shape) while making it easy for the user to view and operate.

[0102] In one embodiment of this specification, the terminal device may display the interaction of the virtual character during video playback (such as the virtual character making corresponding actions or expressions when the video plays to a certain point), or automatically display guidance controls such as "replay" or "view related videos" after the video playback ends, thereby enhancing the immersive experience.

[0103] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps and does not represent the only possible execution order. The order of some steps may be adjusted according to actual needs, or some steps may be omitted. When the claims involve method steps, changes in the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.

[0104] Figure 2 The method described above acquires the user's touch operations on an image displayed on a terminal device, determines the area of ​​the image to which the touch operation is directed, extracts the target element from the area image, and then uses a generative model to generate a virtual avatar with the same characteristics as the target element, along with introductory information about the target element. The virtual avatar and introductory information can then be displayed on the terminal. This allows users to generate corresponding introductory information and virtual avatars for learning and understanding based on simple interactions with images according to their learning needs; it eliminates the need for users to manually search for knowledge related to the target element, while also increasing the fun and personalization of the generated information, enhancing user motivation to learn about the target element, and improving the user experience.

[0105] based on Figure 2In addition to the method described herein, this specification also provides some improved implementation methods, which will be described below.

[0106] In one or more embodiments of this specification, a distinctive virtual avatar can also be generated based on the shooting context information of an image, thereby enhancing the visual appeal of the virtual avatar. Optionally, the method may further include: acquiring the shooting context information of the image, the context information including at least one of shooting location, shooting time, weather conditions, lighting conditions, or surrounding object categories; the step of generating a virtual avatar with the same characteristics as the target element and descriptive information for the target element using a generative model based on the target element may include: generating a prompt word based on the context information and the target element; the prompt word is used to instruct the generative model to adjust the style of the target element in conjunction with the context information to generate a virtual avatar reflecting the context information and descriptive information for the target element; inputting the prompt word into the generative model to obtain the virtual avatar and descriptive information for the target element; wherein the virtual avatar and the descriptive information contain features represented by the context information.

[0107] In one embodiment of this specification, the shooting context information may represent environmental or metadata information related to the image acquisition process, including but not limited to geographical location (such as GPS coordinates), timestamps (such as year, month, day, season, day and night), weather conditions (such as sunny, rainy, snowy), light intensity, and the types of surrounding objects. This information may be obtained through sensors of the terminal device (such as GPS modules, light sensors, gyroscopes) or image analysis algorithms (such as scene classification, object detection).

[0108] In one embodiment of this specification, after inputting prompts into a generative model, the model can extract element features of the target element and context features from the shooting context information. Based on the context features and element features, a virtual avatar with both context features and element features is generated. For example, if the target element is a plant and the context information is autumn, the generated virtual avatar can be depicted as having fallen leaves or changing color; if the context information is a zoo, the virtual animal avatar can wear accessories from a visitor's perspective. This avoids the problem of virtual avatars in object recognition systems being monotonous and detached from real-world contexts, enhancing the user's cognitive association ability. For example, seeing pine cones in autumn in a dry state helps establish a natural concept related to seasonal and plant changes; it also enhances immersion and realism, making educational content closer to life experiences, improving learning interest and memory retention.

[0109] In one or more embodiments of this specification, background music can also be generated for the target element to enrich the display content for the target element and improve the user experience. Optionally, the method may further include: determining feature information of the target element; the feature information includes information representing the type or information representing the material; determining a sound element corresponding to the feature information based on the feature information; performing music art style transfer on the sound element to generate background music; and the terminal device being able to play the background music while playing audio information for the target element.

[0110] In one embodiment of this specification, element features can be extracted from the image using a multimodal model, such as using at least one of the models Material Palette, Qwen 2.5 VL, MobileNet, and ResNet to extract features from elements contained in the image. Material can represent the material category of the element; for example, a seat can include wood, iron, alloy, plastic, etc., or material can also represent physical properties such as the object's hardness, density, elasticity, and structural form. Element category can represent the category to which the element belongs, such as teaching tools, scenery, etc.; or a more granular category, such as tree, grass, fruit, person, vehicle, etc.

[0111] In one embodiment of this specification, the server can first perform element recognition on the region image corresponding to the target element to determine the name of the element contained in the region image, and then determine the material or category corresponding to the element based on a material or category information database. Alternatively, a computer model for identifying materials or categories can be used to identify the material or category of the determined region image to obtain element features such as material information and category information of the elements contained in the region image.

[0112] In one embodiment of this specification, the server can determine the material corresponding to a region image based on color information and texture information contained in the region image, thereby obtaining material information. For example, if the texture is irregular and layered, the edges may have wood pores or a rough feel, the reflective area is large and soft, and the color is brown, then the material can be determined to be wood; if the texture is fine and uniform, some parts have artificial textures such as brushed or frosted finishes, the reflective area is small and bright, and the color is silver, then the material can be determined to be metal.

[0113] In one embodiment of this specification, a sound element can be a sound effect segment for a target element. The server can pre-build a sample library of materials and their corresponding sound elements. The sample library contains original sound samples corresponding to various materials, such as the sound of a pine cone blowing in the wind or a pine cone falling; the sound of a wooden chair corresponding to different actions such as knocking, dragging, or stepping; and the sound of metal products corresponding to different actions such as collision or scratching. Thus, it can select any sound corresponding to a material from the sample library based on the material and determine it as the sound element corresponding to the target element.

[0114] As another implementation method, a material physical parameter library can be pre-established. For example, metals have high density, high hardness, and high sound absorption coefficient; wood has medium density, moderate hardness, and a higher sound absorption coefficient than metal. The server can obtain the corresponding physical parameters from the parameter library based on the material information in the element characteristics; convert the physical parameters into acoustic indicators that can be used for sound adjustment. For example, the fundamental frequency of sound particles can be calculated using density and hardness. The higher the hardness and density, the higher the fundamental frequency is usually; the particle length can be calculated based on the sound absorption coefficient. The smaller the sound absorption coefficient, the slower the sound decays, and the longer the particle length; the vibration of the object can be simulated using the wave equation or the finite element method to determine the basic waveform of the element, and the basic waveform can be adjusted using acoustic indicators to generate the corresponding sound element.

[0115] In one embodiment of this specification, the category can be pre-defined category information used to distinguish object types or categories. The category can include broad categories such as teaching aids, scenery, and animals; it can also include subcategories, such as teaching aids including books, desks, blackboards, and set squares; and scenery including fountains, waterfalls, trees, and landmarks. The object name information represented by the target element can be identified based on the region image; the corresponding category can be determined based on the object name information to obtain the category information of the target element.

[0116] As one implementation method, a sample library of element categories and their corresponding sounds can be pre-built. This sample library can pre-collect the sounds emitted by each element in different actions within each element category. For example, a tree can correspond to the rustling sound of leaves rubbing together, or the sound of leaves colliding. A blackboard can correspond to the sound of friction, or the sound of knocking. After determining the category of the target element, the corresponding sound element can be determined from this sample library.

[0117] As another implementation method, corresponding visual features can be identified based on element categories, and corresponding sound elements can be generated based on these visual features. For example, if the element category is "waterfall," the extracted visual features are rapid water flow and a large drop; the corresponding acoustic parameters can be determined based on these visual features, such as a fundamental frequency concentrated in the mid-low frequency range, high peak amplitude, gentle decay rate, and weak reverberation, generating the sound of water splashing. Similarly, if the element category is "tree," the extracted visual feature is the large swaying of leaves; the corresponding acoustic parameters can be determined based on these visual features, such as a mid-low frequency base superimposed with irregular mid-frequency fluctuations, amplitude varying with the swaying of branches and leaves, and a moderate decay rate, generating the rustling sound of leaves rubbing together. A correspondence between element categories and visual feature types can be pre-established, and corresponding visual features can be extracted based on element categories to obtain corresponding acoustic parameters, and sound elements can be generated based on these acoustic parameters.

[0118] In practical applications, scene recognition can also be performed on images to obtain scene information of target elements. Based on the element features and scene information, the corresponding sound elements can be determined. In this way, the sound elements generated based on element features can conform to the scene information, improving the adaptability of the generated music to the usage scenario.

[0119] In practical applications, preset rules can be used to identify the element features of target elements. For example, the correspondence between the visual features of each element and its material can be preset, thereby determining the corresponding material based on the visual features obtained from the image. Alternatively, neural network models or large models can be used to identify the element features of target elements. Specifically, a model with the function of identifying materials or categories from images can be used to identify target elements and obtain element features. This can be set according to actual needs and is not a specific limitation.

[0120] In one embodiment of this specification, the server can use a sound element as a sound effect segment as background music; or it can perform music style transfer on the sound element to generate background music. Music style transfer can mean transferring the sound characteristics of a certain instrument or the sound characteristics of a certain music genre from the original basic timbre of the element; for example, the basic timbre of a glass is the sound of glass being struck, and music style transfer can mean transferring the harp style from the sound of glass being struck to obtain background music. Music genres can include jazz, rock, country, ballads, etc. In practical applications, artificial intelligence tools can be used to perform music style transfer on sound elements, such as the Artist Brains neural style transfer tool. Artist Brains neural style transfer can extract the texture and spectral characteristics of a sound and blend the sound with another instrument to obtain background music. Other artificial intelligence tools, such as Combobulator, Time Domain Neural Audio Style Transfer, Stability Audio 2.0, etc., can also be used to perform music style transfer on sound elements. The appropriate tool can be selected based on actual needs, and no specific limitation is made here.

[0121] In one embodiment of this specification, the background sound playback control can be the same control as the audio information playback control. When a user operates the playback control, the terminal device can respond by playing the background sound and audio information; alternatively, when a user operates the playback control, the terminal device can respond by pausing the playback of the background sound and audio information. In practical applications, the background sound playback control and the audio information playback control can be different controls, such as a separate background sound playback control and an audio playback control. Users can operate on the corresponding controls to control whether the audio information or background music is played. The terminal device's page can also include a background sound download control, allowing users to download the background sound and store it in the terminal device's local storage space, enabling the user to play the background sound in other applications or pages, thus improving the user experience.

[0122] In one or more embodiments of this specification, the server may further determine ancient poems that match the target element, enabling users to learn about poetry and improving user experience. Optionally, the method may further include: selecting ancient poems that match the target element from a database of ancient poems based on the name and / or category of the target element; the terminal device overlaying the ancient poems onto the background area of ​​the virtual avatar; or, the terminal device displaying the ancient poems in the introductory information.

[0123] In one embodiment of this specification, the server can use keyword retrieval to select ancient poems containing the name or category information of the target element from the ancient poetry database; or the server can select ancient poems that have a mapping relationship with the name or category of the target element; for example, if the target element is a goose, the server can select the ancient poem "Ode to the Goose" from the ancient poetry database and send it to the terminal device so that the terminal device can display it.

[0124] In one embodiment of this specification, the server can also determine the user's educational background based on user characteristics, and display ancient poems corresponding to the target element based on the user's educational background. For example, ancient poems for high school students could be those from high school Chinese textbooks, extracurricular poetry books, or high school exam papers; ancient poems for kindergarten students could be simpler ones, such as "Ode to the Goose." This allows the server to display ancient poems of appropriate difficulty based on the user's educational level, making it easier for the user to learn and use at their current educational stage. In practical applications, users can also select scene tags on the terminal device's page. The server can select ancient poems that match the target element based on the scene corresponding to the scene tag, as well as the name and / or category of the target element. For example, if the target element is a peach tree and the user selects a harvest scene tag, the server can select "Two Poems on Autumn Rain" as the matching ancient poem; if the user selects a blooming flower scene tag, the server can select "Written at the South Village of the Capital" as the matching ancient poem. This ensures that the selected ancient poems match the target element while also meeting the user's scene requirements.

[0125] In one embodiment of this specification, the terminal device overlays classical Chinese poems onto the background area of ​​the virtual avatar. Specifically, the classical Chinese poems can be displayed on a background page, with an opaque or partially transparent display area containing the virtual avatar and introductory information overlaid on the background page. Displaying classical Chinese poems within the introductory information can mean adding classical Chinese poems to the introductory information.

[0126] In one embodiment of this specification, the server can also generate video information matching the classical Chinese poems. During the playback of the video information, the classical Chinese poems are played to enhance the user's interest in learning or watching them. The video information can be an animated video; it can also be multiple images matching the target element and audio information of the classical Chinese poems, etc. The server can also provide introductions based on the background of the classical Chinese poems, expanding the user's knowledge.

[0127] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0128] According to the above explanation, Figure 3 This is a swimlane diagram of a method for generating information based on an image, provided in one embodiment of this specification. For example... Figure 3 As shown, the process includes a user operation stage, a data processing stage, and an information generation stage. This explanation uses the user's interaction with an image displayed on the terminal device's screen as an example. The process may include: The terminal device can display the image that needs to generate information based on the user's touch operation, and performs step 302: display the image.

[0129] Figure 4 This is a schematic diagram of an image displayed in a terminal device according to one embodiment of this specification. Figure 4 As shown, the image can contain elements such as trees, grass, and pine cones.

[0130] Images can be captured using a terminal device; or obtained from the terminal device's local storage; or obtained from the network; or provided by an application, etc.

[0131] Step 304: Obtain the touch operation performed by the user on the image.

[0132] Figure 5 This is a schematic diagram illustrating a user's touch operation on an image, provided as one embodiment of this specification. Figure 5 As shown, virtual actions can be used in the image of the terminal device to represent the user's operation on the pine cone in the image.

[0133] The server can provide user operation information, such as the location of the operation and the displayed image information. The server executes step 306: Determine the area of ​​the image to which the touch operation is pointed.

[0134] In one embodiment of this specification, the server can use the image location of the user's touch operation as the center and expand outwards by a preset distance to obtain a region image containing the target element. Alternatively, the server can call an image segmentation model to extract a region image containing the target element from the image.

[0135] Step 308: Determine the target elements contained in the region image.

[0136] continue Figure 5 The server can determine the target element corresponding to the image location based on the user's touch operation.

[0137] Step 310: Identify at least one characteristic information of the target element, including its name, category, and material.

[0138] In one embodiment of this specification, the server can use a feature recognition model to identify target elements or region images to obtain feature information of the target elements, or it can use a pre-established feature library to determine the feature information of the target elements.

[0139] Step 312: Provide the feature information to the generative model, which can generate a virtual image based on the feature information.

[0140] Step 314: Generate introductory information for the target element.

[0141] In one embodiment of this specification, the generation of introductory information and virtual avatar can be performed in parallel; or the introductory information can be generated first, and then the virtual avatar can be generated; or the virtual avatar can be generated first, and then the introductory information can be generated.

[0142] Step 316: Send the virtual avatar and introductory information to the terminal device.

[0143] After receiving the information sent by the server, the terminal device can perform step 318: display the virtual avatar and introductory information.

[0144] In one embodiment of this specification, the terminal device can render a virtual avatar and introductory information according to a preset template, for example, by jumping from an image page to a new page to display the virtual avatar and introductory information; or it can display the virtual avatar and introductory information as an overlay on the image page. If the virtual avatar and introductory information are displayed as an overlay on the terminal device, the terminal device can have a motion control, allowing the user to move the overlay according to their needs. If the virtual avatar and introductory information are displayed as an overlay on the terminal device, the terminal device can also have a transparency adjustment control, allowing the user to adjust the transparency of the overlay according to their needs. The page displaying the virtual avatar and introductory information can also include download or share controls, so that users can download and share the virtual avatar and introductory information to meet their personalized needs.

[0145] Figure 6 This is a schematic diagram illustrating a virtual image and introductory information, provided as one embodiment of this specification. Figure 6 As shown, a virtual image containing pine cone elements can be displayed, along with introductory information related to pine cones such as "Hi, little friend, I'm a pinecone" and "Pine cones are the seeds of pine trees, which squirrels store for winter." All the introductory information can be accompanied by corresponding audio; or only some of the introductory information can be accompanied by corresponding audio, and so on. This audio can play automatically or based on user interaction. Figure 6The interface may include a play button, such as a speaker-shaped control. Users can click the speaker to play or replay the audio corresponding to the introductory information on the terminal device. In practical applications, different virtual avatars and introductory information can be generated for the same element each time; alternatively, the same virtual avatar and introductory information can be generated for the same element. Besides English, other languages ​​representing the target element's name can also be generated. For example, if the user's frequently used language is Chinese, the introductory information can be displayed in Chinese or played in Chinese audio format. The introductory information may also include information in English or other languages, such as the English name of the target element. If the user's frequently used language is another language besides Chinese, the introductory information can be displayed in the user's frequently used language or played in the user's frequently used language. The introductory information may also include information in Chinese or other languages, such as the Chinese name of the target element, or the corresponding Chinese pinyin, etc. In practical applications, the user's frequently used language is the language information set on the terminal device, or the user can set it in the terminal application.

[0146] In one embodiment of this specification, after the terminal device has finished playing the introductory information, the user can press the play button as follows: Figure 6 The small speaker in the middle is used to operate the terminal device, enabling it to target... Figure 6 Play or replay the English text displayed in the program; or enable the terminal device to target... Figure 6 Play or replay all the introductory information in the video.

[0147] In one embodiment of this specification, some steps are the same as or similar to those in the foregoing embodiments. These foregoing embodiments can be referred to, and will not be repeated here.

[0148] In this way, the server can generate corresponding virtual images and introductory information based on the elements operated by the user. This avoids the situation where the introductory or popular science information for each element is static, thus improving the user experience. It also increases the user's ability to operate on different elements and learn the corresponding introductory information, thereby stimulating the user's learning enthusiasm and interest.

[0149] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.

[0150] Figure 7 For one embodiment of this specification, corresponding to Figure 2 A schematic diagram of the structure of a device based on image generation information.

[0151] like Figure 7As shown, the device may include: The operation acquisition module 702 is used to acquire the touch operation performed by the user on the image displayed on the terminal device; The region image determination module 704 is used to determine the region image pointed to by the touch operation in the image in response to the touch operation; Element determination module 706 is used to extract elements from the region image and determine the target elements contained in the region image; The information generation module 708 is used to generate, based on the target element, a virtual image with the same characteristics as the target element and introductory information about the target element using a generative model; the terminal device is able to display the virtual image and the introductory information.

[0152] based on Figure 7 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.

[0153] Optionally, the device can also be used to: acquire the shooting context information of the image, the context information including at least one of shooting location, shooting time, weather conditions, lighting conditions, or surrounding object categories; the step of generating a virtual image with the same characteristics as the target element and descriptive information for the target element using a generative model based on the target element includes: generating a prompt word based on the context information and the target element; the prompt word is used to instruct the generative model to adjust the style of the target element in combination with the context information to generate a virtual image reflecting the context information and descriptive information for the target element; inputting the prompt word into the generative model to obtain the virtual image and descriptive information for the target element; wherein the virtual image and the descriptive information contain features represented by the context information.

[0154] Optionally, the information generation module can be specifically used to: obtain style tags; the style tags are used to instruct the generative model to generate a virtual image that conforms to the style corresponding to the style tag; the style tags include at least one of cartoon style, 2D style, 3D style, line style, ink painting style, anthropomorphic style, and fairy tale style; provide the target element and the style tags to the generative model, and use the generative model to obtain a virtual image that has the same characteristics as the target element and conforms to the style corresponding to the style tag.

[0155] Optionally, the information generation module can be specifically used to: identify at least one feature information among the name, category, and material of the target element; the step of providing the target element and the style tag to the generative model includes: providing the feature information and the style tag to the generative model, wherein the generative model can generate a virtual image with the same features as the target element and conforming to the style corresponding to the style tag based on the feature information and the style tag.

[0156] Optionally, the information generation module may be specifically used to: acquire user feature information of the terminal device; the user feature information includes the user's historical interaction preference information; the historical interaction preference information represents the interaction information generated by the user interacting with the IP preferred by the user; determine the IP image preferred by the user based on the user feature information; generating a virtual image with the same characteristics as the target element includes: providing the information of the target element and the IP image preferred by the user to a generative model, and using the generative model to obtain a virtual image with the same characteristics as the target element and conforming to the image style of the IP image; the information of the IP image preferred by the user includes at least one of text information described in natural language and an image of the IP image; the text information includes at least one of the name of the IP image, the visual information of the IP image, and the category information of the IP image.

[0157] Optionally, the information generation module may be specifically used to: determine the user's cognitive level based on the user characteristic information; and generate introductory information for the target element that matches the cognitive level.

[0158] Optionally, the information generation module can be specifically used to: generate audio of the introductory information if the cognitive level is lower than a preset level, so that the user terminal can play the audio while displaying the introductory information; or, generate an introductory video based on the introductory information if the cognitive level is lower than a preset level, so that the user terminal can play the introductory information in the form of a video.

[0159] Optionally, the auditory style of the audio has the same emotional dimension as the visual style of the virtual character; or, the auditory and visual style of the introductory video has the same emotional dimension as the visual style of the virtual character; the emotional dimension includes at least one of lively, calm, childlike, realistic, humorous, and serious.

[0160] Optionally, if the introductory information for the target element includes audio information; the audio information is generated using a text-to-speech engine according to the voice that conforms to the image characteristics of the virtual character; the text-to-speech engine can dynamically adjust the timbre, speech rate, and intonation parameters of the audio information according to the image characteristics of the virtual character; the image characteristics include the life stage characteristics, personality characteristics, and physiological attribute characteristics of the virtual character.

[0161] Optionally, if the description information for the target element includes audio information, the page on which the terminal device displays the virtual avatar also includes an audio playback control; if the terminal device receives a user's trigger operation on the audio playback control, then the audio information is played; or, if the description information for the target element includes video information, the page on which the terminal device displays the virtual avatar also includes a video playback control; if the terminal device receives a user's trigger operation on the video playback control, then the video information is played.

[0162] Optionally, the device can also be used to: determine the feature information of the target element; the feature information includes information representing the type or information representing the material; determine the sound element corresponding to the feature information based on the feature information; perform music art style transfer on the sound element to generate background sound; and the terminal device can play the background sound while playing audio information for the target element.

[0163] Optionally, the device can also be used to: select ancient poems that match the target element from the ancient poetry library according to the name and / or category of the target element; display the ancient poems overlaid on the background area of ​​the virtual image; or, display the ancient poems in the introductory information.

[0164] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.

[0165] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0166] The above is an illustrative scheme of an image-based information generation device according to this embodiment. It should be noted that the technical solution of this image-based information generation device and the technical solution of the image-based information generation method described above belong to the same concept. Details not described in detail in the technical solution of the image-based information generation device can be found in the description of the technical solution of the image-based information generation method described above.

[0167] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.

[0168] Figure 8 A structural block diagram of a computing device provided according to one embodiment of this specification is shown.

[0169] The computing device 800 includes: Memory 810 and processor 820; The memory 810 is used to store computer programs / instructions, and the processor 820 is used to execute the computer programs / instructions, which, when executed by the processor 820, implement the steps of the method for generating information based on images.

[0170] Specifically, the components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and the database 850 is used to store data.

[0171] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0172] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0173] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.

[0174] The processor 820 executes the computer instructions to implement the steps of the method for generating information based on images.

[0175] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the image-based information generation method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the image-based information generation method described above.

[0176] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the method for generating information based on an image as described above.

[0177] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the image-based information generation method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the image-based information generation method described above.

[0178] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for generating information based on images.

[0179] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the image-based information generation method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the image-based information generation method described above.

[0180] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the embodiments of apparatus, devices, media, and products, since they are basically similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The apparatus, devices, media, and products provided in the embodiments of this specification correspond to the methods; therefore, the apparatus, devices, media, and products also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, devices, media, and products will not be repeated here.

[0181] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0182] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0183] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0184] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0185] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0186] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0187] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0188] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0189] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0190] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0191] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0192] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0193] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0194] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating information based on an image, comprising: Acquire the touch operations performed by the user in response to the image displayed on the terminal device; In response to the touch operation, determine the region of the image to which the touch operation points in the image; Element extraction is performed on the region image to determine the target elements contained in the region image; Based on the target element, a virtual avatar with the same characteristics as the target element and introductory information about the target element are generated using a generative model; the terminal device is able to display the virtual avatar and the introductory information.

2. The method according to claim 1, further comprising: Obtain the shooting context information of the image, which includes at least one of the following: shooting location, shooting time, weather conditions, lighting conditions, or surrounding object categories; The step of generating a virtual avatar with the same characteristics as the target element and descriptive information about the target element using a generative model, based on the target element, includes: Based on the context information and the target element, prompt words are generated; the prompt words are used to instruct the generative model to adjust the style of the target element in combination with the context information to generate a virtual image that reflects the context information and introductory information for the target element; The prompt words are input into a generative model to obtain the virtual avatar and introductory information for the target element; wherein the virtual avatar and the introductory information contain features represented by the context information.

3. The method according to claim 1, wherein generating a virtual image having the same characteristics as the target element comprises: Get style tags; The style tag is used to instruct the generative model to generate a virtual image that conforms to the style corresponding to the style tag; The style tags include at least one of the following: cartoon style, 2D style, 3D style, line drawing style, ink painting style, anthropomorphic style, and fairy tale style. The target element and the style tag are provided to the generative model, and the generative model is used to obtain a virtual image that has the same characteristics as the target element and conforms to the style corresponding to the style tag.

4. The method according to claim 3, further comprising: Identify at least one feature information among the target element's name, category, and material; Providing the target element and the style tag to the generative model includes: The feature information and the style tag are provided to the generative model, which can generate a virtual image that has the same features as the target element and conforms to the style corresponding to the style tag based on the feature information and the style tag.

5. The method according to claim 1, further comprising: The user feature information of the terminal device is obtained; the user feature information includes the user's historical interaction preference information; The historical interaction preference information represents the interaction information generated when the user interacts with an IP address that the user prefers; Based on the user characteristic information, determine the IP image preferred by the user; The generation of a virtual image having the same characteristics as the target element includes: The target element and the user-preferred IP image information are provided to the generative model, and the generative model is used to obtain a virtual image that has the same characteristics as the target element and conforms to the image style of the IP image; the user-preferred IP image information includes at least one of text information described in natural language and an image of the IP image; the text information includes at least one of the IP image name, the IP image visual information, and the IP image category information.

6. The method according to claim 5, further comprising: Based on the user characteristic information, the user's cognitive level is determined; Generate introductory information for the target element that matches the cognitive level.

7. The method according to claim 6, further comprising: If the cognitive level is lower than the preset level, then the audio of the introductory information is generated so that the user terminal can play the audio while displaying the introductory information; Alternatively, if the cognitive level is lower than a preset level, an introductory video is generated based on the introductory information so that the user terminal can play the introductory information in video format.

8. The method according to claim 7, wherein the auditory style of the audio has the same emotional dimension as the visual style of the virtual character; or, the auditory and visual style of the introductory video has the same emotional dimension as the visual style of the virtual character; the emotional dimension includes at least one of lively, calm, childlike, realistic, humorous, and serious.

9. The method according to claim 1, wherein the introductory information for the target element includes audio information; the audio information is audio information generated by a text-to-speech engine according to the voice that conforms to the image characteristics of the virtual image; the text-to-speech engine can dynamically adjust the timbre, speech rate and intonation parameters of the audio information according to the image characteristics of the virtual image; the image characteristics include the life stage characteristics, personality characteristics and physiological attribute characteristics of the virtual image.

10. The method according to claim 1, wherein if the introductory information for the target element includes audio information, the page on which the terminal device displays the virtual image further includes an audio playback control; if the terminal device receives a user's trigger operation on the audio playback control, then the audio information is played; or, If the description information for the target element includes video information, the page on which the terminal device displays the virtual avatar also includes a video playback control; if the terminal device receives a user's trigger operation on the video playback control, then the video information is played.

11. The method according to claim 1, further comprising: Determine the feature information of the target element; The feature information includes information indicating the type or information indicating the material; Based on the feature information, determine the sound element corresponding to the feature information; The sound element is subjected to musical art style transfer to generate background sound; the terminal device can play the background sound while playing audio information for the target element.

12. The method according to claim 1, further comprising: Based on the name and / or category of the target element, select ancient poems that match the target element from the ancient poetry database; The terminal device overlays the ancient poems onto the background area of ​​the virtual avatar; or, the terminal device displays the ancient poems in the introductory information.

13. An apparatus for generating information based on images, comprising: The operation acquisition module is used to acquire the touch operations performed by the user on the image displayed on the terminal device; A region image determination module is used to determine the region image pointed to by the touch operation in the image in response to the touch operation; An element determination module is used to extract elements from the region image and determine the target elements contained in the region image. The information generation module is used to generate, based on the target element, a virtual image with the same characteristics as the target element and introductory information about the target element using a generative model; The terminal device can display the virtual avatar and the introductory information.

14. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.