Potential image generation method and device, computer storage medium and electronic equipment
Patent Information
- Application Number
- CN202311096074.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-08-28
AI Technical Summary
但是上述定制处理的实现成本高且可选的图片数量有限,不能根据终端设备的状态自适应产生符合要求的潜在图像,例如自拍照无法根据设备特性动态显示
[0040]As can be seen from the above technical solution, in this application, behavioral text vectors representing device behavior states and attribute text vectors representing device display attributes are obtained; image text vectors and image feature elements of the base image are obtained; various vector values of the behavioral text vectors, attribute text vectors, and image text vectors are combined to obtain related vector combinations. These related vector combinations represent the combination relationship between various image feature elements and device behavior states and attributes, and subsequently, the requirements for potential images under various behavior states are determined based on these combinations. Next, all related vector combinations are input into a pre-trained semantic script generation model for processing to generate semantic script text describing the requirements for generating potential images; then, the semantic script text and image feature elements are input into a pre-trained conditional image generation model for processing, modifying and reorganizing the image feature elements according to the requirements described in the semantic script text to generate potential images corresponding to the device behavior states; finally, the corresponding potential images are displayed based on the device behavior states. In this way, through the above processing, potential images suitable for the current terminal state can be adaptively generated according to different states of the terminal device and the base image.
Smart Images

Figure CN117152286B_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, and more particularly to a method, apparatus, computer storage medium, and electronic device for generating a potential image. Background Technology
[0002] With the rapid development of electronic technology, electronic devices have appeared in various forms in our lives and work. Examples include mobile terminal devices with flip or folding functions, and television terminals that support screen rotation. Unlike traditional display devices, these new devices have external characteristics such as rotation and folding.
[0003] In traditional devices, display size and viewing methods are relatively fixed, and the display state of images on the screen is also relatively fixed. For example, the display size, orientation, and content of a wallpaper remain unchanged. However, for newer devices with features such as rotation and folding, although the device state changes, image display still uses existing methods. For instance, if a user sets a photo as wallpaper on a folding device, the wallpaper will be cut and displayed on the cover screen after the device is folded. Figure 1 As shown, this wallpaper display method is clearly not user-friendly.
[0004] Therefore, for new devices with features such as rotation and folding, users have new demands: they want the images displayed on the screen to change accordingly with the device's state to adapt to the updated device status. For example, for foldable terminals, users want the wallpaper images displayed to change depending on whether the screen is open or folded.
[0005] Currently, to meet the above requirements, it is necessary to customize the terminal devices, creating specific wallpaper themes for various terminal states, such as... Figure 2a The wallpaper shown is an open flower when the screen is folded up and a closed flower when the screen is folded down, or, as... Figure 2b The screen shown folds left and right, with butterfly wings fluttering as the device opens and closes. However, the aforementioned customized processing is costly to implement and has a limited number of selectable images. It cannot adaptively generate suitable potential images based on the terminal device's state; for example, selfies cannot be dynamically displayed according to device characteristics. Summary of the Invention
[0006] This application provides a method, apparatus, computer storage medium, and electronic device for generating latent images, which can adaptively generate latent images suitable for the current terminal state based on different states of the terminal device and the base image.
[0007] To achieve the above objectives, this application adopts the following technical solution:
[0008] A method for generating and displaying a potential image, comprising:
[0009] Obtain the behavior text vector used to characterize the device's behavioral state and the attribute text vector used to characterize the device's display attributes;
[0010] Obtain the image text vector and image feature elements of the base image;
[0011] Various vector values of the behavioral text vector, the attribute text vector, and the image text vector are combined to obtain a related vector combination; wherein, the related vector combination is a combination consisting of a value of one behavioral text vector, a value of one attribute text vector, and a value of one image text vector;
[0012] All the relevant vectors are combined and input into a pre-trained semantic script generation model for processing to generate semantic script text that describes potential image generation needs;
[0013] The semantic script text and the image feature elements are input into a pre-trained conditional image generation model for processing to generate a potential image corresponding to the device's behavioral state.
[0014] The corresponding potential image is displayed based on the behavior state of the device.
[0015] Preferably, in the semantic script generation model, each of the input related vector combinations is processed by a long short-term memory network, and valid processing results are identified; attention weights corresponding to each of the related vector combinations are generated based on an attention mechanism; and the attention weights are used to weight and fuse each of the valid processing results to obtain the semantic script text.
[0016] Preferably, the identification of valid processing results includes: if there is no difference in the processing results at different times corresponding to the relevant vector combination, then it is determined to be an invalid processing result.
[0017] Preferably, the semantic script generation model is a hierarchical attention neural network based on long short-term memory.
[0018] Preferably, the conditional image generation model is a latent diffusion model. In the conditional image generation model, a target image is generated based on the image feature elements, and the semantic script text and the image feature elements are fused to obtain a multimodal vector. Target parameters for the behavioral state are generated based on the multimodal vector, and the generated target image is verified and adjusted based on the target parameters to obtain the latent image for the behavioral state.
[0019] The results generated at different times are compressed into the latent feature space to learn the representation of the latent image.
[0020] Preferably, in the conditional image generation model, the generated latent image preferentially changes the domain-related image feature elements indicated by the semantic script text, while preferentially keeping the image feature elements that are not related to the domain unchanged.
[0021] Preferably, the target parameters in the behavioral state include the target parameters corresponding to different time states of the behavioral state;
[0022] The latent image of the behavioral state includes the latent image corresponding to each of the different time states of the behavioral state.
[0023] There are multiple potential images corresponding to the device's behavioral state, including potential images corresponding to different time points.
[0024] Preferably, obtaining the behavioral text vectors used to characterize different behavioral states of the device includes:
[0025] The device detects different behavioral states and generates corresponding behavioral text vectors.
[0026] Preferably, obtaining the relevant vector combination includes:
[0027] The various combinations of vector values of the behavioral text vector, the attribute text vector, and the image text vector are subjected to relevance filtering to obtain multiple combinations of vector values with a relevance greater than a set threshold, which are used as the relevant vector combinations; wherein, each combination of vector values is a combination consisting of a value of the behavioral text vector, a value of the attribute text vector, and a value of the image text vector.
[0028] Preferably, the relevance filtering of various vector value combinations of the behavioral text vector, the attribute text vector, and the image text vector includes:
[0029] For each vector value combination, calculate the PMI value between each pair of word vector values in the combination, and then calculate the sum of the PMI values as the correlation of the vector value combination.
[0030] A potential image generation and display apparatus, comprising: a device state and attribute acquisition unit, a basic image processing unit, a filtering unit, a semantic script generation model processing unit, a conditional image generation model processing unit, and a display unit;
[0031] The device status and attribute acquisition unit is used to acquire behavioral text vectors that characterize different behavioral states of the device and attribute text vectors that characterize different display attributes of the device.
[0032] The basic image processing unit is used to obtain the image text vector and image feature elements of the basic image;
[0033] The filtering unit is used to combine various vector values of the behavior text vector, the attribute text vector, and the image text vector to obtain a related vector combination; wherein, the related vector combination includes a combination consisting of a value of the behavior text vector, a value of the attribute text vector, and a value of the image text vector;
[0034] The semantic script generation model processing unit is used to combine all the relevant vectors and input them into a pre-trained semantic script generation model for processing, thereby generating semantic script text that describes potential image generation needs.
[0035] The conditional image generation model processing unit is used to input the semantic script text and the image feature elements into a pre-trained conditional image generation model for processing, and generate potential images corresponding to different behavioral states of the device.
[0036] The display unit is used to display the corresponding potential image based on the behavior state of the device.
[0037] A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, can implement the method for generating and displaying a potential image as described in any of the preceding claims.
[0038] An electronic device, comprising at least a computer-readable storage medium and a processor;
[0039] The processor is configured to read executable instructions from the computer-readable storage medium and execute the instructions to implement the method for generating and displaying the potential image as described in any of the preceding embodiments.
[0040] As can be seen from the above technical solution, in this application, behavioral text vectors representing device behavior states and attribute text vectors representing device display attributes are obtained; image text vectors and image feature elements of the base image are obtained; various vector values of the behavioral text vectors, attribute text vectors, and image text vectors are combined to obtain related vector combinations. These related vector combinations represent the combination relationship between various image feature elements and device behavior states and attributes, and subsequently, the requirements for potential images under various behavior states are determined based on these combinations. Next, all related vector combinations are input into a pre-trained semantic script generation model for processing to generate semantic script text describing the requirements for generating potential images; then, the semantic script text and image feature elements are input into a pre-trained conditional image generation model for processing, modifying and reorganizing the image feature elements according to the requirements described in the semantic script text to generate potential images corresponding to the device behavior states; finally, the corresponding potential images are displayed based on the device behavior states. In this way, through the above processing, potential images suitable for the current terminal state can be adaptively generated according to different states of the terminal device and the base image. Attached Figure Description
[0041] Figure 1 An illustration of the existing fixed wallpaper display;
[0042] Figure 2a An illustration of existing custom wallpapers Figure 1 ;
[0043] Figure 2b Illustration 2 for existing custom wallpapers;
[0044] Figure 3 This is a schematic diagram illustrating the basic process of the potential image generation and display method in this application;
[0045] Figure 4a and Figure 4b Examples of the base image and its corresponding potential image in this application;
[0046] Figure 5 This is a schematic diagram illustrating the specific process of the potential image generation and display method in the embodiments of this application;
[0047] Figure 6 For Figure 4a and Figure 4b The diagram shows a flowchart of a latent image generation and display method using the base image and its corresponding latent image as an example.
[0048] Figure 7 This is a schematic diagram of the filter processing in a specific embodiment of this application;
[0049] Figure 8This is a structural example diagram of the semantic script generation model in a specific embodiment of this application;
[0050] Figure 9 This is a structural example diagram of the conditional image generation model in a specific embodiment of this application;
[0051] Figure 10 Example diagram of the potential image generation process;
[0052] Figure 11 An example of potential image generation and display using the method of this application. Figure 1 ;
[0053] Figure 12 Figure 2 is an example of potential image generation and display using the method of this application;
[0054] Figure 13 An example of potential image generation and display using the method of this application. Figure 3 ;
[0055] Figure 14 This is a schematic diagram of the basic structure of the potential image generation and display device in this application;
[0056] Figure 15 This is a schematic diagram of the basic structure of the electronic device provided in this application. Detailed Implementation
[0057] To make the objectives, technical means, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings.
[0058] The basic idea of this application is to combine the device's behavior state and attribute information, utilize the text vectors and image elements of the base image, and adaptively generate potential images that conform to different behavior states and attributes of the device through a neural network model.
[0059] Figure 3 This is a schematic diagram illustrating the basic flow of the method for generating and displaying potential images in this application. For example... Figure 3 As shown, the method includes:
[0060] Step 301: Obtain behavioral text vectors that characterize different behavioral states of the device and attribute text vectors that characterize different display attributes of the device.
[0061] To achieve the function of adaptively generating latent images based on device behavior states, it is first necessary to obtain the different behavior states of the device. In this application, text vectors are used to represent the different behavior states of the device, and these text vectors are called behavior text vectors B. At the same time, it is also necessary to obtain text vectors representing different display attributes of the device, and these text vectors are called attribute text vectors X, such as different screen sizes of the device, or the rotation angle and rotation speed of the device's display screen.
[0062] Step 302: Obtain the image text vector and image feature elements of the base image.
[0063] In this application, a user-defined base image is analyzed to obtain image feature elements and a text vector (hereinafter referred to as the image text vector) used to describe the image. The image feature elements can be implemented using methods such as image segmentation and recognition. The image text vector can also be obtained using various existing methods, such as analyzing the results of image segmentation and recognition.
[0064] The processes in steps 301 and 302 can be performed in parallel or in any order.
[0065] Step 303: Combine the various vector values of the behavioral text vector, attribute text vector, and image text vector to obtain the relevant vector combination.
[0066] Different values of behavioral text vectors represent different behavioral states of the device, different values of attribute text vectors represent different display attributes of the device, and different values of image text vectors can represent different image feature elements of the base image. In this application, each combination of text vector values is a combination of a behavioral text vector value, an attribute text vector value, and an image text vector value. Through different combinations of these three types of text vector values, combinations of different device behavioral states, different display attributes, and different image feature elements are represented.
[0067] In this step, each combination of values for the three types of text vectors is determined as a related vector combination, which is used to represent various cases of the three types of text vector combinations.
[0068] Alternatively, considering computational limitations, when determining the relevant vector combinations in this step, the correlation of each value combination of the three types of text vectors can be determined. Then, based on the correlation results, filtering can be performed to obtain value combinations with correlation exceeding a set threshold. These are used as relevant vector combinations to represent various cases where the three types of text vectors are strongly correlated. In this case, the vector combinations represent the strong correlation between various image feature elements and device behavior states and attributes. Subsequently, based on these strong correlation combinations, the demand for potential images under various behavior states can be determined.
[0069] Step 304: Combine all relevant vectors and input them into the pre-trained semantic script generation model for processing to generate semantic script text that describes potential image generation needs.
[0070] This step utilizes a neural network model to generate semantic script text based on all relevant vector combinations. The semantic script text describes the generation requirements of the latent image, that is, what requirements the latent image should meet. These latent image generation requirements correspond to different behavioral states and attribute combinations of the device, guiding the generation of latent images corresponding to these different behavioral states.
[0071] The neural network model used in this step is a pre-trained semantic script generation model. Optionally, this model can be a hierarchical attention neural network based on long short-term memory, which has two significant features:
[0072] a. It can reflect the hierarchical structure of words and sentences.
[0073] b. Attention mechanism is a long-term memory mechanism that can intuitively give the contribution of each word or the result of a sentence. It has two levels of attention mechanism applied at the word and sentence levels, which enables it to pay attention to increasingly important content when constructing semantic script representations.
[0074] More specifically, in the semantic script generation model, each relevant vector combination in the input is processed by a Long Short-Term Memory (LSTM) network, and valid processing results are identified. Attention weights are generated for each relevant vector combination based on an attention mechanism. These attention weights are then used to weight and fuse the valid processing results to obtain the semantic script text. Optionally, when identifying valid processing results, if the processing results for a corresponding relevant vector combination are indistinguishable at different times, they can be determined as invalid processing results.
[0075] In this way, on the one hand, the Long Short-Term Memory network can determine the time-related feature attributes for each relevant vector combination, specifically the feature states under long and short time periods; on the other hand, the hierarchical attention mechanism can be used to configure appropriate attention weights for each relevant vector combination, so that when generating semantic script text, attention is focused more on the more important relevant vector combinations, that is, the generation requirements of potential images are reflected more on the more important relevant vector combinations.
[0076] Step 305: Input the semantic script text and image feature elements into the pre-trained conditional image generation model for processing to generate a potential image corresponding to the device behavior state.
[0077] This step utilizes a neural network model to generate potential images that adapt to different behavioral states of the device.
[0078] Specifically, the neural network model used in this step is a pre-trained conditional image generation model, mainly comprising a latent diffusion model. The latent diffusion model, as a generative model, aims to describe the basic data distribution and estimate model parameters by minimizing the difference between the actual and generated data distributions. In this application, the latent diffusion model is used for processing, and the results of each step are compressed into a high-quality latent feature space to learn latent representations. The image feature elements of the base image are modified and reorganized according to the requirements of the semantic script text description to generate latent images corresponding to the device behavior states. Specifically, in the conditional image generation model, a target image is generated based on image feature elements, and the semantic script text and image feature elements are fused to obtain a multimodal vector. Target parameters for the device behavior state are generated based on the multimodal vector, and the generated target image is then verified and adjusted based on the target parameters to obtain a latent image for the device behavior state that meets the verification requirements. The target parameters for the device behavior state can be a set, in which case the corresponding latent image for that device behavior state can be a single image, which can be displayed as a static image; or, the target parameters for the device behavior state can be multiple sets corresponding to different time states, in which case the corresponding latent images for that device behavior state can also be multiple images corresponding to different time states, which can be displayed as dynamic images.
[0079] Furthermore, in conditional image generation models, the generated latent image can preferentially modify domain-relevant image feature elements indicated by the semantic script text, while preferentially keeping image feature elements unrelated to the corresponding domain unchanged. For example... Figure 4a and 4b As shown, Figure 4a The base image's feature elements include the child's head, arms, body, speech bubbles, background trees, and colors. Figure 4b For the generated multiple potential images, such as Figure 4b As shown, when generating these latent images, the child's arms gradually close. The semantic script text can be used to analyze the domain-related image features, including the child's head and arms. When generating latent images, these image features are modified first. The domain-independent image features include the background forest, body, and color. When generating latent images, these image features are left unchanged.
[0080] Step 306: Display the corresponding potential image based on the device's behavior status.
[0081] Through the aforementioned steps, a potential image corresponding to the device's behavioral state can be generated. This step displays the corresponding potential image based on the current behavioral state of the device.
[0082] The potential image generated earlier, corresponding to a device behavior state, can be a single image or multiple potential images arranged in chronological order. These images can be displayed as video images, such as live wallpapers.
[0083] This concludes the process for generating and displaying the latent image in this application. As can be seen from the above, this application can combine device state and display attributes, modify and reorganize the image feature elements of the base image, and adaptively generate a latent image that matches the device state and display attributes.
[0084] The specific implementation of this application is described below through specific embodiments. Figure 5 This is a schematic diagram illustrating the specific process of the potential image generation and display method in the embodiments of this application. Figure 6 This is a schematic diagram illustrating the flowchart of a latent image generation and display method, using the base image shown in Figure 4 as an example. Figure 5 and Figure 6 As shown, the method in this embodiment specifically includes:
[0085] Step 501: Obtain the device's resource information.
[0086] Users can enable the corresponding function before potential image generation is needed. For example, a user can set a new wallpaper on a foldable device and turn on the AI wallpaper generation function.
[0087] Step 502: Detect the behavior state of the device and generate behavior text vector B, and obtain the attribute text vector X of the device.
[0088] In this embodiment, a corresponding potential image is generated and displayed by real-time detection of the device's current behavioral state. Of course, in other specific implementations, various behavioral states of the device can be obtained in advance through various device operations, and corresponding potential images can be generated and saved. When the device performs an operation and enters a new behavioral state, the saved potential image corresponding to the new behavioral state will be displayed.
[0089] In this embodiment, it is assumed that the behavior is a phone folding operation, corresponding to the generation of behavior text vector B1, such as folding, standing, unfolding, etc.; the device's display attributes are obtained: the main display size and the cover display size, which are 1812*2176 and 904*2316 respectively, corresponding to attribute text vectors X1 and X2. This step can be performed as follows: Figure 6 Implemented in the device behavior discriminator shown.
[0090] Step 503: Obtain the image feature elements E and the image text vector I of the base image.
[0091] In this embodiment, the processing in this step is implemented through a visual encoder and decoder model, specifically... Figure 6 This is achieved through the image recognition component. The visual encoder segments and recognizes the base image to obtain image feature elements, and then uses a decoder (e.g., an autoregressive language model) to generate image-text vectors corresponding to these feature elements. For example... Figure 6 As shown, the acquired image feature elements include image E1 representing the background quantity, image E2 representing the child quantity, and image E3 representing the bubbles. The acquired image text vectors include I1, I2, and I3. Here, I1 represents the background forest, I2 represents the child with outstretched arms, and I3 represents three bubbles.
[0092] The processes in steps 502 and 503 are executed after step 501 and before step 504. Steps 502 and 503 can be executed in parallel or in any order.
[0093] Step 504: Filter the various combinations of values for the behavioral text vector, attribute text vector, and image text vector to obtain relevant vector combinations.
[0094] In this step, all acquired behavioral text vectors, attribute text vectors, and image text vectors are input into the filter, such as... Figure 7 As shown, the filter combines all inputs with different values. Each combination includes three text vector values, each belonging to a different text vector category. For each combination, the relevance of the three text vector values is determined. Combinations with a relevance higher than a set threshold are selected from all combinations, i.e., combinations with high relevance are filtered out. In specific implementation, the filter can calculate the relevance or trend convergence based on the point mutual information (PMI) between word vectors in the word vector combination. The larger the PMI sum, the stronger the relevance. Based on the set threshold, the corresponding word vectors are filtered. The PMI value is the sum of the PMI values between any two word vectors in the word vector combination; PMI1(word1, word2) = P(word1 & word2). P is a word relevance input model preset based on device attributes. This model collects and trains the convergence distribution of various things under known changes in device attributes. In this application, the word vector combination in the filter is the value of the three text vectors.
[0095] Step 505: Combine all relevant vectors and input them into the pre-trained semantic script generation model for processing to generate semantic script text that describes potential image generation needs.
[0096] As mentioned earlier, the semantic script generation model is a hierarchical attention neural network based on LSTM. The neural network's parameters are pre-trained using training samples to generate the semantic script generation model. In this step, this semantic script generation model is used to process all relevant vector combinations of the input to generate semantic script text.
[0097] This embodiment provides a structural example of a semantic script generation model, SSG, such as... Figure 8 As shown in the diagram, all relevant vector combinations are input into the SSG model, which comprises multiple identical branch models, each processing one relevant vector combination. Within each branch model, the input relevant vector combinations are first processed by an embedding layer to obtain embedded features. Then, they are processed by LSTM at different time levels to reflect the state or degree of change of the embedded features over time. The LSTM results from each time level are merged, and dropout is used to determine the validity of the processing results, identifying valid results (dropout is typically used during the training phase to prevent overfitting and reduce error). Results with insignificant changes over time are considered invalid, while others are considered valid. Simultaneously, an attention mechanism is used to obtain attention weights for each relevant vector combination. The valid processing results and their corresponding attention weights are weighted and fused to ultimately obtain the generation requirements for the potential image describing the device's behavioral state, represented using text, i.e., semantic script text.
[0098] In addition, during the training of the semantic script generation model, softmax is used to compare the predicted semantic script text with the actual script text to obtain the loss function result. Based on this result, the various parameters of the model are adjusted in reverse until the model training completion condition is met.
[0099] Step 506: Input the semantic script text and image feature elements into the pre-trained conditional image generation model for processing to generate a potential image corresponding to the device's behavioral state.
[0100] As mentioned earlier, the conditional image generation model mainly includes the latent diffusion model, which is also a neural network model. The neural network's various parameters are trained in advance using training samples to generate the conditional image generation model. In this step, this conditional image generation model is used to modify and recombine the input image feature elements according to the latent image generation requirements expressed in the input semantic script text, generating a latent image corresponding to the device's behavioral state.
[0101] This embodiment provides a structural example of a conditional image generation model, ITTI, such as... Figure 9As shown in the diagram, the model comprises a diffusion processing module, a conditional input layer, and a backpropagation denoising module. Its specific functions include: the diffusion processing module, used for image generation, which continuously adjusts pixel values in the generated image based on the partial differential equations in the backpropagation denoising module, gradually bringing them closer to the target image; the conditional input layer, used to fuse text features from the semantic script with image features to obtain a multimodal vector, and generates target parameters (T(y)) for a specific state based on the predicted semantic script; and the backpropagation denoising module, used to describe the propagation process of matter in the image, calculating the verification and adjustment of the image generated by the diffusion processing module in the current state.
[0102] In addition, in the conditional image generation model, the domain-specific features given in the semantic footstep text (such as specific parts like the head and arms) are modified first (i.e., modified as much as possible), while domain-independent features (such as background, color, and body) are preserved.
[0103] The latent image generated by the conditional image generation model corresponds to a specific device behavior state, which in this embodiment corresponds to B1. Furthermore, the latent image corresponding to the device behavior state may be a single image, or it may be multiple latent images arranged in chronological order, as shown in Figure 4. When generating multiple latent images arranged in chronological order, the generation process of these latent images can... Figure 10 The formula in the text represents, where X1 to X T The state space represents the transition process from one state to another, where y represents the input conditional data, T(y) represents the posterior probability obtained after adding conditional data at each time step t, α(t) represents the correlation coefficient of the current state during the process, and the higher the coefficient, the closer it is to the target image. σ 2 p(x) is the hyperparameter of the Gaussian distribution variance, N represents the Markov chain computation process, and p(x) is the hyperparameter of the Gaussian distribution variance. t |y) represents the output of a certain state in the state space. Figure 10 The image on the left is the original image, and the image on the right is X. T The target image generated at each time step.
[0104] Step 507: Display the generated potential image based on the device's behavioral state.
[0105] Display the corresponding generated potential image based on the device's current behavior state, for example Figure 6 The image shown is of a child with their arms folded.
[0106] This concludes the method flow in this embodiment.
[0107] The above embodiments provide a specific example of generating a latent image. In fact, the method of this application can also adaptively generate corresponding latent images for various device behavior states. For example... Figure 11 As shown, for the base image on the left, after detecting the behavior of the device folding to the right, multiple potential images arranged in chronological order can be generated and displayed using the method described above, visually appearing as the main target in the image gradually moving to the right; after detecting the behavior of the device folding to the left, multiple potential images arranged in chronological order can be generated and displayed using the method described above, visually appearing as the main target in the image gradually moving to the left. Or, for example... Figure 12 As shown, for the base image on the left, the phone's folding action is detected. The latent image generated using the above method visually appears as a small boat moving up and down with the folding action. Alternatively, for example... Figure 13 As shown, for static wallpapers set on rotating televisions, the method described in this application can extract the feature elements of the base image, combine them with the rotation attributes of the device, and generate multiple potential images to form an animated wallpaper, with the ball rolling as the screen rotates.
[0108] The above describes the specific implementation of the latent image generation and display method in this application. Through the method described above, a matching latent image can be adaptively generated and displayed by combining device behavior state and attribute information. This satisfies the user's need for image changes as device state changes, eliminates the need to customize images based on the terminal, enriches the user's image choices, and provides a better user experience.
[0109] This application also provides a potential image generation and display apparatus that can be used to implement the method described above. Figure 14 This is a schematic diagram of the basic structure of the device. Figure 14 As shown, the device includes: a device status and attribute acquisition unit, a basic image processing unit, a filtering unit, a semantic script generation model processing unit, a conditional image generation model processing unit, and a display unit.
[0110] The device status and attribute acquisition unit is used to acquire behavioral text vectors that characterize different behavioral states of the device and attribute text vectors that characterize different display attributes of the device.
[0111] The basic image processing unit is used to obtain the image text vector and image feature elements of the basic image;
[0112] The filtering unit combines various vector values from behavioral text vectors, attribute text vectors, and image text vectors to obtain a relevant vector combination. This relevant vector combination consists of a value from one behavioral text vector, a value from one attribute text vector, and a value from one image text vector.
[0113] The semantic script generation model processing unit is used to combine all the relevant vectors and input them into a pre-trained semantic script generation model for processing, thereby generating semantic script text that describes potential image generation needs.
[0114] The conditional image generation model processing unit is used to input semantic script text and image feature elements into a pre-trained conditional image generation model for processing, and generate potential images corresponding to different behavioral states of the device.
[0115] The display unit is used to display the corresponding potential image based on the device's behavior state.
[0116] This application also provides a computer-readable storage medium that stores instructions, which, when executed by a processor, can perform the steps in the method for generating and displaying a potential image as described above. In practical applications, the computer-readable medium may be included in the devices / apparatus / systems of the above embodiments, or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium stores instructions, which, when executed by a processor, can perform the steps in the method for generating and displaying a potential image as described above.
[0117] According to the embodiments disclosed in this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof, but not intended to limit the scope of protection of this application. In the embodiments disclosed in this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0118] Figure 15 An electronic device is also provided for this application. For example... Figure 15 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:
[0119] The electronic device may include a processor 1501 with one or more processing cores, a memory 1502 with one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When the program in the memory 1502 is executed, a method for generating and displaying a potential image can be implemented.
[0120] Specifically, in practical applications, this electronic device may also include components such as a power supply 1503 and an input / output unit 1504. Those skilled in the art will understand that... Figure 15 The structure of the electronic device shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0121] The processor 1501 is the control center of the electronic device. It connects various parts of the electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1502, and calling data stored in the memory 1502, it performs various functions of the server and processes data, thereby monitoring the electronic device as a whole.
[0122] Memory 1502 can be used to store software programs and modules, i.e., the aforementioned computer-readable storage medium. Processor 1501 executes various functional applications and data processing by running the software programs and modules stored in memory 1502. Memory 1502 may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the server, etc. In addition, memory 1502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 1502 may also include a memory controller to provide processor 1501 with access to memory 1502.
[0123] The electronic device also includes a power supply 1503 that supplies power to the various components. This power supply can be logically connected to the processor 1501 via a power management system, thereby enabling functions such as charging, discharging, and power consumption management. The power supply 1503 may also include one or more DC or AC power supplies, a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator, or any other components.
[0124] The electronic device may also include an input / output unit 1504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, and optical signal inputs related to user settings and function control. The input unit output 1504 can also be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which can be composed of graphics, text, icons, video, and any combination thereof.
[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating and displaying a potential image, comprising: Obtain the behavior text vector used to characterize the device's behavioral state and the attribute text vector used to characterize the device's display attributes; Obtain the image text vector and image feature elements of the base image; Various vector values of the behavioral text vector, the attribute text vector, and the image text vector are combined to obtain a related vector combination; wherein, the related vector combination is a combination consisting of a value of one behavioral text vector, a value of one attribute text vector, and a value of one image text vector; All the relevant vectors are combined and input into a pre-trained semantic script generation model for processing to generate semantic script text that describes potential image generation needs; The semantic script text and the image feature elements are input into a pre-trained conditional image generation model for processing to generate a potential image corresponding to the device's behavioral state. The corresponding potential image is displayed based on the behavior state of the device.
2. The method according to claim 1, characterized in that, In the semantic script generation model, each of the input related vector combinations is processed by a long short-term memory network, and valid processing results are identified; attention weights corresponding to each of the related vector combinations are generated based on an attention mechanism. The semantic script text is obtained by weighting and fusing the effective processing results using the attention weights.
3. The method according to claim 2, characterized in that, The identification of valid processing results includes: if there is no difference in the processing results at different times corresponding to the relevant vector combination, then it is determined to be an invalid processing result.
4. The method according to claim 1 or 2, characterized in that, The semantic script generation model is a hierarchical attention neural network based on long short-term memory.
5. The method according to claim 1, characterized in that, The conditional image generation model is a latent diffusion model. In the conditional image generation model, a target image is generated based on the image feature elements, and the semantic script text and the image feature elements are fused to obtain a multimodal vector. Target parameters for the behavioral state are generated based on the multimodal vector, and the generated target image is then verified and adjusted based on the target parameters to obtain the latent image for the behavioral state. The results generated at different times are compressed into the latent feature space to learn the representation of the latent image.
6. The method according to claim 1, characterized in that, In the conditional image generation model, the generated latent image preferentially changes the domain-related image feature elements indicated by the semantic script text, while preferentially keeping the image feature elements that are not related to the domain unchanged.
7. The method according to claim 5, characterized in that, The target parameters under the behavioral state include the target parameters corresponding to different time states under the behavioral state; The latent image under the behavioral state includes the latent image corresponding to each of the different time states under the behavioral state; There are multiple potential images corresponding to the device's behavioral state, including potential images corresponding to different time points.
8. The method according to claim 1, characterized in that, The acquisition of behavioral text vectors used to characterize different behavioral states of the device includes: The device detects different behavioral states and generates corresponding behavioral text vectors.
9. The method according to claim 1, characterized in that, The process of obtaining the relevant vector combination includes: The various combinations of vector values of the behavioral text vector, the attribute text vector, and the image text vector are subjected to relevance filtering to obtain multiple combinations of vector values with a relevance greater than a set threshold, which are used as the relevant vector combinations; wherein, each combination of vector values is a combination consisting of a value of the behavioral text vector, a value of the attribute text vector, and a value of the image text vector.
10. The method according to claim 9, characterized in that, The step of performing relevance filtering on various combinations of vector values of the behavioral text vector, the attribute text vector, and the image text vector includes: For each vector value combination, calculate the PMI value between each pair of word vector values in the combination, and then calculate the sum of the PMI values as the correlation of the vector value combination.
11. A device for generating and displaying a potential image, characterized in that, The device includes: a device status and attribute acquisition unit, a basic image processing unit, a filtering unit, a semantic script generation model processing unit, a conditional image generation model processing unit, and a display unit; The device status and attribute acquisition unit is used to acquire behavioral text vectors that characterize different behavioral states of the device and attribute text vectors that characterize different display attributes of the device. The basic image processing unit is used to obtain the image text vector and image feature elements of the basic image; The filtering unit is used to combine various vector values of the behavior text vector, the attribute text vector, and the image text vector to obtain a related vector combination; wherein, the related vector combination includes a combination consisting of a value of the behavior text vector, a value of the attribute text vector, and a value of the image text vector; The semantic script generation model processing unit is used to combine all the relevant vectors and input them into a pre-trained semantic script generation model for processing, thereby generating semantic script text that describes potential image generation needs. The conditional image generation model processing unit is used to input the semantic script text and the image feature elements into a pre-trained conditional image generation model for processing, and generate potential images corresponding to different behavioral states of the device. The display unit is used to display the corresponding potential image based on the behavior state of the device.
12. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the instructions are executed by the processor, they can implement the method for generating and displaying the potential image as described in any one of claims 1 to 10.
13. An electronic device, characterized in that, The electronic device includes at least a computer-readable storage medium and a processor; The processor is configured to read executable instructions from the computer-readable storage medium and execute the instructions to implement the method for generating and displaying a potential image as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Text recognition method and system, electronic equipment and storage medium
CN114943960A
Phrase-level text image generation method and system based on self-attention mechanism
CN115587160A