UI stylization method and device based on extended reality

By generating stylized description instructions and determining the stylized target UI elements, the problem of mismatch in prompt information in real-life devices is solved, and the user experience continuity in immersive scenarios and the effect of natural integration of UI into the environment is achieved.

CN120014210AActive Publication Date: 2025-05-16BEIJING IRISVIEW TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510481099.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

When a user enters an immersive scene using the expansion reality device, the prompt information from outside the scene is interrupted due to style mismatch, and the user experience needs to be stylized to match the current scene.

Method used

By obtaining the expanded real-life scenarios and UI elements to be stylized in the user's vision, using the large language model to generate stylized description instructions, integrating the stylized UI elements into the visual expression of the expanded real-life scenario, and determining the stylized target UI elements through the text image model, and rearranging to generate stylized UI components.

Benefits of technology

It realizes real-time adjustment of the visual performance of the UI in an immersive scenario, avoids the abruptness of prompt information, ensures the continuity of the user experience, and allows the UI to naturally integrate into various immersive environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014210A_ABST
    Figure CN120014210A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a UI stylization method and device based on expanded reality, and the method comprises the steps: obtaining a scene in the view of a user and a to-be-stylized UI element; according to the scene and the to-be-stylized UI element, a stylized description instruction for the to-be-stylized UI element is generated, and the stylized description instruction is used for fusing the to-be-stylized UI element into visual expression of an expanded reality scene; according to the stylized description instruction, determining each single target UI element in the stylized description instruction; and rearranging each single target UI element to generate a stylized UI component. According to the scheme, the visual performance of the UI to be stylized is adjusted in real time according to the specific style of the XR scene, and the experience of a user in an immersive scene is not interrupted. Through machine learning and deep learning technologies, a training model identifies and adapts to different scene styles, so that the UI can be naturally fused into a user immersion environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more particularly to a method and device for stylizing a UI based on augmented reality. Background Art

[0002] When a user uses an extended reality (XR) device to enter an immersive scene (for example, when playing an XR game), they still need to receive prompt information outside the scene where the user is, such as important information prompts, device battery message prompts, and prompts from other applications. However, the prompt information outside the scene where the user is located only has the original default style, which is somewhat abrupt in the immersive scene where the user is currently located. In particular, strong reminder prompts may even interrupt the user's experience. Therefore, it is necessary to stylize the prompt information outside the scene where the user is located into the scene where the user is currently located to achieve the stylization of the prompt information.

[0003] Traditional UI stylization solutions rely on SDK interfaces, which results in separation of functionality and visual style and insufficient stylization. Summary of the invention

[0004] In order to solve the problem of experience gap caused by the mismatch between prompt information and immersive application style in the prior art, an embodiment of the present specification provides a UI stylization method and device based on augmented reality, the method comprising: obtaining an augmented reality scene and a UI element to be stylized in a user's field of view; generating a stylized description instruction for the UI element to be stylized according to the augmented reality scene and the UI element to be stylized, the stylized description instruction being used to integrate the UI element to be stylized into the visual expression of the augmented reality scene; determining each individual stylized target UI element in the stylized description instruction according to the stylized description instruction; and rearranging each individual stylized target UI element to generate a stylized UI component.

[0005] According to one aspect of an embodiment of the present specification, generating a stylized description instruction for the UI element to be stylized based on the augmented reality scene and the UI element to be stylized includes: inputting the visual expression corresponding to the augmented reality scene and the UI element to be stylized into a large language model, and obtaining the stylized description instruction output by the large language model; wherein the large language model is obtained through pre-training.

[0006] According to one aspect of an embodiment of the present specification, each individual stylized target UI element in the stylized description instruction is determined according to the stylized description instruction, including: inputting the stylized description instruction into a text image model to obtain a single stylized target UI element output by the text image model; the text image model is obtained by training an initial text image model through a stylized sample description instruction.

[0007] According to one aspect of an embodiment of the present specification, each individual stylized target UI element in the stylized description instruction is determined according to the stylized description instruction, including: converting the stylized description instruction into a corresponding code; determining the corresponding UI element in the stylized description instruction according to the code, and obtaining multiple stylized individual stylized target UI elements.

[0008] According to one aspect of an embodiment of the present specification, the text image model is trained in the following manner: obtaining a training sample set, the training sample set including: sample text and a label UI element corresponding to the sample text; inputting the sample text into an initial text image model to obtain a single predicted UI element output by the initial text image model; constructing a loss function based on the label UI element and the predicted UI element, and iteratively updating the parameters of the initial text image model based on the loss function until the loss function converges to a preset threshold, thereby constructing a text image model.

[0009] According to one aspect of an embodiment of the present specification, after a text image model is constructed and a single stylized target UI element is output by the text image model, the method further includes: determining a block image and text loss function based on image features of each target UI element and text features corresponding to the target UI element; determining a visual feature extraction loss function based on visual features of the target image and visual features of a scene image in the visual expression, wherein the target image is obtained by combining all target UI elements; determining a scene consistency loss function based on features of the target image and a scene description of the scene image; and guiding fine-tuning of at least one of a large language model and a text image model based on at least one of the block image and text loss function, the visual feature extraction loss function, and the scene consistency loss function.

[0010] According to one aspect of the embodiments of the present specification, the large language model further outputs an arrangement layout for the scene image and the UI elements to be stylized, and the rearranging of each individual stylized target UI element to generate a stylized UI component includes: according to the scene image and the arrangement layout of the UI to be stylized, the rearranging of each individual stylized target UI element to generate a stylized UI component.

[0011] According to one aspect of the embodiments of this specification, the method further includes:

[0012] The total loss function is constructed using the following formula:

[0013] + + ;

[0014] in, , , They represent the weights of the block image and text loss function, the scene consistency loss function, and the normalized visual feature extraction loss function respectively; represents the block image and text loss function, represents the scene consistency loss function, Represents the normalized visual feature extraction loss function.

[0015] An embodiment of the present specification provides a UI stylization device based on augmented reality, the device comprising: an acquisition unit, used to acquire an augmented reality scene and a UI element to be stylized in a user's field of view; a first generation unit, used to generate a stylized description instruction for the UI element to be stylized based on the augmented reality scene and the UI element to be stylized, the stylized description instruction being used to integrate the UI element to be stylized into the visual expression of the augmented reality scene; a determination unit, used to determine each single target UI element in the stylized description instruction based on the stylized description instruction; and a second generation unit, used to rearrange each single target UI element to generate a stylized UI component.

[0016] An embodiment of the present specification also provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the UI stylization method based on augmented reality when executing the computer program.

[0017] The embodiments of the present specification also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for stylizing the UI based on augmented reality is implemented.

[0018] This solution adjusts the visual performance of the UI in real time according to the specific style of the XR scene without interrupting the user's experience in the immersive scene. Through machine learning and deep learning technology, the model is trained to recognize and adapt to different visual styles, allowing the UI to naturally integrate into various immersive environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 The figure is a flow chart of a UI stylization method based on augmented reality according to an embodiment of the present specification;

[0021] Figure 2 The figure is a flow chart of a method for determining a single target UI element according to an embodiment of the present specification;

[0022] Figure 3 The figure shows a schematic diagram of a stylized UI effect according to an embodiment of this specification;

[0023] Figure 4 The figure is a flow chart of a method for training a text image model according to an embodiment of the present specification;

[0024] Figure 5 The figure is a flow chart of a method for constructing a loss function to perform model fine-tuning according to an embodiment of the present specification;

[0025] Figure 6 The figure shows a schematic diagram of a process of performing UI stylization based on an augmented reality scene according to an embodiment of the present specification;

[0026] Figure 7 The figure shows a structural schematic diagram of a UI stylization device based on augmented reality according to an embodiment of the present specification;

[0027] Figure 8 The figure shows a schematic diagram of a UI to be stylized according to an embodiment of the present specification;

[0028] Fig. 9 The figure shows a schematic diagram of a scene image of a user playing a game according to an embodiment of the present specification;

[0029] Fig.10 The figure shows a schematic diagram of the structure of a computer device according to an embodiment of the present specification.

[0030] Description of the accompanying symbols:

[0031] 701, acquisition unit;

[0032] 702, a first generating unit;

[0033] 703. Determine the unit;

[0034] 704, second generation unit;

[0035] 1002. Computer equipment;

[0036] 1004. Processor;

[0037] 1006. Memory;

[0038] 1008. Driving mechanism;

[0039] 1010, input / output module;

[0040] 1012. Input device;

[0041] 1014. Output device;

[0042] 1016. Presentation equipment;

[0043] 1018. Graphical user interface;

[0044] 1020. Network interface;

[0045] 1022. Communication link;

[0046] 1024. Communication bus. DETAILED DESCRIPTION

[0047] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0048] It should be noted that the terms "first", "second", etc. in the description and claims of this specification and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of this specification described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0049] This specification provides method operation steps as described in the embodiments or flow charts, but more or fewer operation steps may be included based on routine or non-creative work. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the system or device product is executed in practice, it can be executed in the order of the method shown in the embodiments or the drawings or in parallel.

[0050] It should be noted that the UI stylization method and device based on augmented reality in this specification can be used in the field of image processing and in the field of natural language processing. This specification does not limit the application field of the UI stylization method and device based on augmented reality.

[0051] Figure 1 The flowchart of a UI stylization method based on augmented reality in an embodiment of this specification is shown, which specifically includes the following steps:

[0052] Step 101: Obtain an augmented reality scene and UI elements to be stylized in the user's field of view.

[0053] In this specification, the scene in the user's field of view is obtained based on the application or software currently used and operated by the user using the extended reality device (such as an XR device). In this specification, the scene in the user's field of view may include a variety of different visual expressions, such as: scene images, scene texts, 3D objects, etc. Among them, it can be achieved by recording a video presented in the user's field of view when the user uses the XR device, and capturing one or more frames in the video as the scene image; or capturing the scene image in the user's field of view when the user uses the XR device at a preset time point; further, it is also possible to capture the scene image in the user's field of view when the user uses the XR device at that time point at a random time point.

[0054] In the embodiments of this specification, the scene in the user's field of view is related to the application, software or tool currently used by the user in the XR device. Specifically, when the user uses the XR device to play a game, the scene that appears in the user's field of view is an image related to the game scene; when the user uses the social software in the XR device to chat, the scene that appears in the user's field of view is an image related to the chat dialog box; when the user uses the news website in the XR device to check economic information, the scene that appears in the user's field of view is an image related to the news report.

[0055] In the embodiments of the present specification, the UI elements to be stylized include: a UI to be stylized for prompting the user or having a notification function, a UI to be stylized for executing an operation function, a UI to be stylized with a specific function, etc. Among them, the content of the UI elements to be stylized includes a variety of UI elements, and the UI elements to be stylized usually have a system default style. If the content of the UI to be stylized with the system default style is directly pushed to the user who is currently in the immersive scene, the prompts from outside the scene (for example, from other applications) are inevitably abrupt, and the content of the UI to be stylized does not match the style of the immersive application. In particular, strong reminders may even interrupt the user's experience. Therefore, in the present specification, it is necessary to convert the UI elements to be stylized into new stylized prompt UI elements that are similar in style to the scene in the user's field of view, so as to ensure that the new stylized UI elements are naturally integrated into the visual style of the scene in the user's field of view.

[0056] like Figure 8 Shown is a schematic diagram of a UI to be stylized in an embodiment of this specification. The figure shows the UI to be stylized described in this specification. In the upper left corner of the figure, there is a faded title "Exercise Reminder" and a purple close button ("X") in the upper right corner. In the upper middle of the UI to be stylized, there is a reminder message: "Don't forget to do some exercise today!", with a weightlifting emoticon next to it. There are two interactive buttons below the reminder message: "Do it!" (black background and white text, emphasizing immediate action); "Remind me later" (gray background and black text, providing the option of reminding later).

[0057] Step 102: Generate a stylized description instruction for the UI element to be stylized according to the augmented reality scene and the UI element to be stylized, wherein the stylized description instruction is used to integrate the UI element to be stylized into the visual expression of the augmented reality scene.

[0058] In this specification, the system uses a large language model to identify the scene style obtained in step 101 and the UI elements to be stylized, and generates a stylized description instruction that matches the scene style, so as to use the stylized description instruction to stylize the UI elements to be stylized. The stylized instruction is used to integrate the UI elements to be stylized into the scene image, and the stylized description instruction includes: the requirements for designing a new UI for the visual expression in the augmented reality scene in the user's current field of view, and the requirements need to include: each UI element (including: UI shape, line, color, text and brightness, etc.) and the layout form and position of each UI element in the scene image.

[0059] For example, when a user uses an XR device to play a science fiction-style XR game and receives a UI element to be stylized, this step needs to convert the style of the UI element to be stylized according to the science fiction style of the XR game.

[0060] Specifically, the visual expression corresponding to the scene in step 101 and the UI element to be stylized are input into the large language model to obtain the stylized description instruction output by the large language model; wherein the large language model is obtained by pre-training. In this specification, the scene image in the visual expression corresponding to the scene and the UI to be stylized can be simultaneously input into the large language model to obtain the stylized description instruction.

[0061] In this specification, the stylized description instruction includes stylizing each element of the UI and generating a detailed text description. The stylized description instruction is performed in the system and the user is not aware of it.

[0062] like Figure 8 As shown, large language model recognition Figure 8 The schematic diagram of the UI element to be stylized is shown. After the large language model identifies the UI element to be stylized, the converted and generated text content is as follows:

[0063] a. Title and close button:

[0064] In the upper left corner of the UI to be stylized, there is a faded title "Exercise Reminder".

[0065] In the upper right corner of the UI to be stylized, there is a purple close button ("X"), which the user can click to close the pop-up window.

[0066] b. Reminder information:

[0067] In the upper center of the UI to be stylized, there is a reminder message: "Don’t forget to do some exercise today!", with a weightlifting emoji next to it, adding fun and visual appeal.

[0068] c. Interactive buttons:

[0069] The user can select two buttons:

[0070] “Do it!” The button has a black background and white text, emphasizing immediate action.

[0071] Remind me later, a button with black text on a grey background, provides the option to be reminded later.

[0072] This manual generates stylized description instructions through a large language model, and generates a new stylized UI without being restricted by a predefined structure, ensuring that UI components are naturally integrated into the visual style of the user scenario in an immersive environment while retaining their core functionality and interactivity.

[0073] Fig. 9 The diagram shows a scene image of a user playing a game in an embodiment of the present specification, which includes an augmented reality scene in the user's field of vision corresponding to step 101. After the large language model recognizes the content and style of the immersive game scene and interprets the scene style, the converted generated text content is as follows:

[0074] The style of this VR scene has a strong sense of futuristic technology. Here are some of its features:

[0075] Use of colors: Blue and purple neon lights are mainly used to create a high-tech and mysterious atmosphere.

[0076] Geometric Design: There are many circular and rectangular geometric patterns in the scene, giving it a sophisticated and modern feel.

[0077] Lighting effect: The use of luminous rings and lines enhances the three-dimensional and technological feel of the scene.

[0078] Materials and Textures: The surface looks smooth and metallic, further emphasizing the futuristic design.

[0079] Overall, the scene feels like something out of a sci-fi movie or a high-tech laboratory.

[0080] According to the above description, the content and style obtained by the large language model respectively identifying the game scene played by the user, and the text content generated after the large language model identifies the UI to be stylized are combined to generate a stylized description instruction for the game scene, as shown below:

[0081] Design a modern, high-tech user interface (UI) popup that includes the following elements:

[0082] Title: Use a futuristic font and a gradient blue to display “Sports Reminder”.

[0083] Close button: Design a glowing purple “X” button that users can click to close the popup.

[0084] Reminders: A reminder message that reads, "Don't forget to do some sports today!" with a tech-style weightlifting emoji next to it to convey vitality and energy.

[0085] Interaction Buttons:

[0086] “Go For it!”: Designed with a black background and glowing white text, the button has a subtle blue glow around the edge to emphasize immediate action.

[0087] “Remind me”: Designed with a grey background and black text, with a faint glow around the edge of the button to provide the option to remind you later.

[0088] Background: Use a dark, tech-style background with abstract lines and shapes to enhance the futuristic feel. Make sure the overall design matches the style of the modern VR scene, and use blue and purple neon lighting effects to create a high-tech and mysterious atmosphere.

[0089] Step 103: Determine, according to the stylized description instruction, each single stylized target UI element in the stylized description instruction.

[0090] In this specification, the stylized description instruction may be a stylized description instruction in text form, or a corresponding code form.

[0091] In some embodiments of the present specification, if the stylized description instruction is in the form of natural language text, determining each individual target UI element in the stylized description instruction according to the stylized description instruction includes: inputting a stylized description image in the form of natural language text into a text image model (for example, a diffusion model), and obtaining a single target UI element output by the text image model that corresponds to the stylized description; wherein the text image model is obtained by training an initial text image model using the stylized sample description instruction.

[0092] In some embodiments of this specification, if the stylized description instruction is in code form, the stylized description instruction is converted into code form, and a certain UI element among the UI elements to be stylized corresponding to the stylized description instruction is determined according to the code. Further, in the process of inputting the stylized description instruction into the text image model, the stylized description instruction can also be converted into code, and the code can be input into the text image model for prediction. For details, see Figure 2 describe.

[0093] In this specification, the text image model is a multimodal model that can process images and texts, and is used for matching between images and texts. It has functions such as text parsing, UI layout generation, and image rendering. In the UI stylization scenario, the text image model can be used for style transfer. Specifically, a known UI element to be stylized is matched with a certain style using the text image model. For example, an ordinary UI design is compared with a "future" style or a "hand-painted" style to optimize the image appearance of the UI design.

[0094] In this specification, any one of the CLIP model, ALIGN model, and any one of the StableDiffusion model or GANs can be used to build a text image generation model for stylized UI. For details about the training process of the text image model, see Figure 4 describe.

[0095] Step 104: rearrange the individual stylized target UI elements to generate a stylized UI component.

[0096] The multiple single target UI elements obtained according to step 103 are not arranged in order. In this step, all the single target UI elements can be rearranged and arranged according to the style of the augmented reality scene in the user's field of view and the arrangement and layout of the UI elements to be stylized to generate an overall stylized UI component, such as Figure 3 As shown. In one embodiment of the present specification, each individual target UI element is rearranged according to a preset arrangement layout of the system to generate a stylized UI component. The preset arrangement layout includes: a preset arrangement standard, a preset interface, or a preset arrangement layout template. These preset arrangement layouts can be automatically generated by relying on a large language model. In some other embodiments of the present specification, individual target UI elements can also be rearranged according to the arrangement layout of UI elements of historical scenes that have been pre-processed and are similar to the augmented reality scenes in the user's field of view. This step does not limit the basis for rearranging the target UI elements.

[0097] This step redefines the stylized target UI element based on the single stylized target UI element, so that each stylized target UI element can be naturally integrated into the immersive environment of the user's game scene.

[0098] Figure 2 The figure is a flow chart of a method for determining a single target UI element according to an embodiment of the present specification. Figure 2 The process of converting the stylized description instructions into code and inputting them into the text image model for prediction is shown, which specifically includes the following steps:

[0099] Step 201: convert the stylized description instruction into corresponding code.

[0100] According to the text in the style description instruction content, generate the modification code corresponding to the style of the UI element in the UI code to be stylized. Specifically, input the stylized description instruction into a model with instruction conversion function to obtain the code corresponding to the stylized description instruction. Wherein, for each sentence of text in the stylized description instruction content, generate the corresponding code; or for each paragraph of text in the stylized description instruction content, generate the corresponding code. Wherein, the style of UI elements includes but is not limited to: UI styles in different interaction states, sound feedback, etc.

[0101] Step 202, determine the corresponding UI element in the stylized description instruction according to the code, and obtain a stylized single target UI element. In this step, a new stylized UI element is rendered according to the modified code, and a single stylized target UI element is generated. In this specification, a single target UI element includes but is not limited to: stylized buttons, icons, and text boxes. The generated single target UI elements all conform to the style attributes of the augmented reality scene.

[0102] Figure 4 The flowchart of a method for training a text image model according to an embodiment of the present specification is shown, which specifically includes the following steps:

[0103] Step 401: Acquire a training sample set, wherein the training sample set includes: sample text and a label UI element corresponding to the sample text.

[0104] In order to build a text image model, this manual prepares a training sample set in advance. Among them, the sample text in the training sample set can be data in the form of code. Specifically, the code can be a complete code or multiple sub-codes, and sample codes of various types and styles can be collected from open source projects, code repositories or online programming platforms. When the sample code is a complete code, the code corresponds to multiple UI elements. When the sample code is multiple sub-codes, each sub-code corresponds to a UI element. Furthermore, the training sample set also includes a label UI element corresponding to each sample code. The label UI element is used to indicate that the sample code describes the appearance attributes of a button or an icon, such as color, material, transparency, etc.

[0105] Step 402: input the sample text into an initial text image model to obtain a single predicted UI element output by the initial text image model.

[0106] According to the above steps, the initial text image can be a CLIP model, an ALIGN model, etc. In this step, before inputting the sample text into the initial text image model, the sample text needs to be converted into an input format acceptable to the model, such as word embedding or code vector. Furthermore, the output layer of the initial text image model is designed to generate pixel values ​​or feature vectors corresponding to UI elements.

[0107] Based on the input sample text, the initial text image model outputs multiple single predicted UI elements. Before the initial text image model is iterated and fine-tuned multiple times, the single predicted UI elements output by the initial text image model have a large gap in style and visual effects from the extended reality scene and are not accurate enough. Therefore, it is necessary to further adjust the parameters and parameter weights in the initial text image model.

[0108] Step 403, constructing a loss function according to the label UI element and the predicted UI element, iteratively updating the parameters of the initial text image model based on the loss function until the loss function converges to a preset threshold, and constructing a text image model.

[0109] In this step, the difference between the predicted UI elements and the labeled UI elements generated by the model is evaluated by selecting a suitable loss function. Specifically, the cross entropy loss function can be used to calculate the loss value of the predicted UI elements and the labeled UI elements, and the optimizer is used to iteratively update the parameters and weights of the initial text image model according to the loss value to minimize the loss function until the loss function converges to a preset threshold, and it is considered that the text image model is constructed.

[0110] Figure 5 The flowchart of a method for constructing a loss function to fine-tune a model in an embodiment of this specification is shown. In this specification, after the text image model is constructed, each model can be fine-tuned to improve the closeness between the UI elements predicted by the model and the extended reality scene. This method specifically includes the following steps:

[0111] Step 501 : determining a block image-text loss function according to image features of each target UI element and text features corresponding to the target UI element.

[0112] In this specification, since a complete UI component is usually composed of discrete elements (such as buttons, backgrounds, etc.), different UI elements among multiple target UI elements are prioritized by modifying the block image and text loss in the target UI element, so as to more finely control some UI elements (for example, only highlighting one button), so that different styles can be applied to different parts of the UI elements.

[0113] In this specification, the formula of the block image loss function is as follows:

[0114] ;

[0115] in, represents the block image and text loss function, The feature extractor representing the text image model, whose output is a feature vector; represents the i-th target UI element output by the text image model, Represents the image features of the i-th target UI element; Indicates the text or code text corresponding to the i-th target UI element. Represents the text feature corresponding to the i-th target UI element, Indicates the weight of the i-th target UI element. In the block image loss function, the image features of the target UI element and the text features corresponding to the i-th target UI element are extracted through the feature extractor of the text image model. The block image loss function extracts the image features of the target UI element and the features corresponding to the corresponding text description through the feature extractor, and measures the difference between the two. Cosine similarity Used to measure the similarity between image features and text features. The similarity is converted into a difference metric. The block image-text loss function reflects the difference between the image features of the target UI element and the features corresponding to the corresponding text description. Used to adjust the contribution of different UI elements to the total loss.

[0116] Step 502: Determine a visual feature extraction loss function based on the visual features of the target image and the visual features of the scene image in the visual representation.

[0117] In this step, the visual features (such as color and texture) in the scene image in the field visual expression are extracted through the feature extraction network or visual encoder and transferred to the target image, and then the difference between the visual features in the scene image and the visual features in the target image is calculated to form a visual feature loss function. , the target image is recorded as The target image is an image composed of all single target UI elements obtained after all the aforementioned steps are processed.

[0118] The formula of the visual feature extraction loss function is as follows:

[0119] ( encoder ( ;

[0120] ;

[0121] in, represents the visual feature extraction loss function, Represents the visual encoder function, which is used to extract high-level visual features from the input image. It can represent any model that can effectively extract image features, such as convolutional neural network (CNN), Vision Transformer (ViT), etc. Represents the visual features of the scene image, Represents the visual features of an image composed of all individual target UI elements; It can represent the observed The maximum value or pre-set value; Represents the normalized visual feature extraction loss function. In this specification, the visual feature extraction loss function is intended to reflect the difference between the overall visual features of an image composed of all individual target UI elements and the overall visual features of a scene image. After determining the value of the visual feature extraction loss function based on the visual features of the UI prediction image and the visual features of the scene image, determine whether the value of the loss function meets a preset threshold, and if not, continue to fine-tune the text image model and the large language model.

[0122] Step 503: Determine a scene consistency loss function according to the features of the target image and the scene description of the scene image.

[0123] Specifically, the formula for scene consistency loss is as follows:

[0124] ), ( ;

[0125] in, represents the scene consistency loss function, represents a feature extractor, whose output is a feature vector; A text description representing a scene image, represents the text features extracted from the scene image, represents the entire target image, Represents the image features extracted from the target image. Cosine similarity ), ( Used to measure the similarity between text features and image features. Convert similarity to a difference measure.

[0126] In this step, the text features extracted from the scene image are used to obtain the text description related to the overall style or overall theme of the scene image, such as "soft forest theme" or "modern technology sense" and the like; the image features extracted from the complete target image. In this specification, the scene consistency loss function is intended to reflect the difference between the image features of the target image and the text features of the scene image.

[0127] The scene consistency loss function can make the final generated target image visually coherent with the original scene image after stylization, so that the stylized target image matches the scene description of the scene image.

[0128] In the embodiment of this specification, the method further includes:

[0129] The total loss function is constructed using the following formula:

[0130] + + ;

[0131] in, , , They represent the weights of the block image and text loss function, the scene consistency loss function, and the normalized visual feature extraction loss function respectively; represents the block image and text loss function, represents the scene consistency loss function, represents the normalized visual feature extraction loss function. The specific values ​​of the three weights in the formula can be tuned through experiments.

[0132] In this specification, the scene stylization requirements may change, so dynamic weights are introduced in the loss function to adjust the influence of each loss term in the loss function in real time. For example, when the scene visual features are more distinct, the weight of the visual feature extraction loss can be increased, and when the scene style is simpler, the weight of the content loss can be increased to maintain the clarity of the UI.

[0133] The formula for dynamic weight is as follows:

[0134] = =g ;

[0135] in, represents a function that adjusts weights according to scene complexity, =g Represents a function that adjusts weights based on scene clarity.

[0136] Step 504: guiding at least one of the large language model and the text image model to perform fine-tuning according to at least one of the block image and text loss function, the scene consistency loss function, and the normalized visual feature extraction loss function.

[0137] In this specification, fine-tuning the large language model using a loss function includes: fine-tuning part of or all of the large language model using a loss function.

[0138] When the amount of data is large enough, all layers of the large language model are fine-tuned; when the amount of data is small, the top layer of the large language model is fine-tuned to reduce the consumption of computing resources while still allowing the model to have sufficient flexibility to learn new tasks. In the embodiments of this specification, the bottom layer of the large language model can also be fixed, and only the top layer of the large language model is fine-tuned, keeping the bottom feature representation learned by the large language model during the training phase unchanged, while enabling the top layer to better adapt to new tasks.

[0139] In some other embodiments of the present specification, the large language model can also be fine-tuned layer by layer starting from the bottom layer through a loss function until all layers of the large language model are fine-tuned, focusing on using continuous prompts (such as embedding vectors) to adjust the behavior of the large language model rather than directly modifying the model's weights.

[0140] In some embodiments of the present specification, fine-tuning the text image model includes: adjusting the code input to the text image model, so as to obtain a UI element that is more in line with the scene image style effect, and then generating an image. Through the scene consistency loss function, the constructed large language model and text image model are iteratively fine-tuned until the scene consistency loss function value converges to a preset threshold, thereby determining that the large language model and text image model have completed fine-tuning.

[0141] In some embodiments of the present specification, fine-tuning of a large language model using a visual feature extraction loss function includes: freezing the parameters of the large language model, and only optimizing a series of continuous task-specific vectors (i.e., prefixes) to achieve the optimization task. The lightweight design avoids the waste of storage and computing resources while maintaining the performance of the large language model.

[0142] In the embodiments of this specification, a low-rank adapter method can be used to fine-tune the large language model. For example, the low-rank adaptation (LoRA) method is used to freeze the weights of the large language model, and the trainable rank decomposition matrix is ​​injected into each layer of the Transformer architecture, significantly reducing the number of trainable parameters in downstream tasks while maintaining model quality performance. This specification can also use the dynamic low-rank adaptation (DyLoRA, Dynamic Low-RankAdaptation) method to introduce upper projection and lower projection matrices, as well as a mechanism for dynamically adjusting the rank, in response to the problems of the LoRA block (such as fixed size and difficulty in rank optimization). The training speed is faster, the performance is almost unchanged, and excellent performance is shown over a wider range of ranks. In this specification, the fine-tuned large language model generates stylized description instructions that are more intelligent and close to user usage habits based on the scene image and the UI to be stylized.

[0143] For example, according to Figure 8 In the embodiment of the embodiment, the stylized description instruction output by the large language model may be: "Show a reminder message in the middle: "Don't forget to do some exercise today!", add a technology-style weightlifting emoji next to it to convey vitality and energy; and add a crossbow symbol from the game to fit the user's current game scene." Among them, adding the weightlifting emoji and the crossbow symbol to the stylized description instruction are both intelligent results output by the large language model after fine-tuning.

[0144] In the embodiments of this specification, in addition to outputting stylized description instructions for scene images and UI to be stylized, the large language model can also output the arrangement and layout of scene images and UI to be stylized through the model's own learning. Among them, the arrangement and layout of the augmented reality scene and UI elements to be stylized include: determining the area where the UI to be stylized can be laid out in the scene image, the position of each element in the UI to be stylized in the UI to be stylized, the size of the scene image, etc.

[0145] For example, determine that the close button ("X") in the UI to be stylized is in the upper left corner or upper right corner of the scene image, determine that the text in the UI to be stylized should be in the blank space of the scene image, or can be overlaid on the current scene image, the UI elements in the UI to be stylized should be located within the scene image area, but should not exceed the size of the scene image, the reminder message: "Don't forget to do some exercise today!" is located above the interactive buttons provided to the user, and the interactive buttons are set in the middle position or below the middle position of the system UI.

[0146] Therefore, after determining the arrangement layout, the multiple single UI elements obtained in the previous text are arranged according to the area of ​​the scene image and the layout and arrangement order in the UI to be stylized, so as to obtain a complete stylized image, such as Figure 3 shown.

[0147] Figure 3 The figure shows a schematic diagram of a stylized UI effect of an embodiment of the present specification. In the upper left corner of the figure, a futuristic font is used to display "Sports Reminder" in a gradient blue. A glowing purple "X" button is set in the upper right corner, and the user can click to close the pop-up window. A reminder message is displayed in the middle of the figure: "Don't forget to do some sports today! Don't forget to do some sports today!" Next to it is a technological weightlifting emoticon to convey vitality and energy. In the figure, interactive buttons are set in the lower third of the page: "Go For it! Go do it!" and "Remind me to remind me later." Figure 3A dark, tech-style background is used with abstract lines and shapes to enhance the futuristic feel. The overall design matches the style of a modern VR scene, using blue and purple neon lighting effects to create a high-tech and mysterious atmosphere.

[0148] In this specification, before obtaining the scene in the user's field of view, the UI element to be stylized, and generating the stylization instruction, the method further includes:

[0149] According to the scene in the user's field of view, it is determined whether the scene is a multi-task scene. In this step, by intercepting the image in the user's field of view, it is determined whether the scene image in the user's field of view involves multiple scenes. Specifically, the scene images in the user's field of view can be continuously collected within a certain time period to determine whether multiple scene images of different styles appear within the time period.

[0150] If yes, use the default notification UI style. In this step, when the scene is a multi-task scene, it is not necessary to stylize the UI to be stylized according to different task scenes.

[0151] If not, the scene image is stylized. In this step, when the scene is a single-task scene, the stylized UI needs to be stylized according to the style of the scene, so as to perform Figure 1 to Figure 2 and Figures 4 to 5 Steps shown.

[0152] Figure 6 The figure is a schematic diagram of a process of performing UI stylization based on an augmented reality scenario according to an embodiment of the present specification.

[0153] Figure 6 In the process, the ocean-style XR scene content and the UI to be stylized in the user's field of view are first obtained (the UI to be stylized is represented by A in the first sub-image in the figure). The XR scene in the user's field of view shows the visual expression of marine life, including images of jellyfish and 3D images of sharks, as well as visual expression elements such as light effects, tones, and textures in immersive ocean environment scenes. The ocean-style XR scene content and the UI to be stylized (A) are input into the large language model, and the text in the UI to be stylized (A) and the image description in the XR scene are output, and they are combined to generate stylized description instructions.

[0154] Figure 6In the second sub-image, the stylized description instruction obtained in the first image is further converted into corresponding codes, and the codes are input into the text image model to obtain the stylized single target UI element in the stylized description instruction determined by the text image model. Then, according to the arrangement and layout of various visual expressions in the XR scene, the stylized target UI elements are rearranged according to the corresponding arrangement and layout to obtain the arranged stylized UI components shown in the third sub-image, and the visual properties of multiple single target UI elements are adjusted as a whole to obtain the stylized UI components shown in the fourth sub-image. Such stylized UI components are immersively deployed in the XR scene without interrupting the user's immersive experience in the current scene.

[0155] like Figure 7 The figure shows a schematic diagram of the structure of a UI stylization device based on mixed reality in an embodiment of the present specification. In this figure, the basic structure of the UI stylization device based on mixed reality is described, wherein the functional units and modules can be implemented in software, or a general chip or a specific chip can be used to implement UI stylization based on mixed reality. The device specifically includes:

[0156] An acquisition unit 701 is used to acquire the scene in the user's field of view and the UI element to be stylized;

[0157] A first generating unit 702 is used to generate a stylized description instruction for the UI element to be stylized according to the scene and the UI element to be stylized, wherein the stylized description instruction is used to integrate the UI element to be stylized into the scene;

[0158] A determination unit 703, configured to determine, according to the stylized description instruction, each single target UI element in the stylized description instruction;

[0159] The second generating unit 704 is used to rearrange the individual target UI elements to generate a stylized UI component.

[0160] like Fig.10As shown, it is a schematic diagram of a computer device provided by an embodiment of this specification. The UI stylization method based on augmented reality described in this application can be applied to the computer device. The computer device 1002 may include one or more processors 1004, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 1002 may also include any memory 1006, which is used to store any kind of information such as code, settings, data, etc. Non-restrictive, for example, the memory 1006 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory device, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 1002. In one case, when the processor 1004 executes an associated instruction stored in any memory or a combination of memories, the computer device 1002 can perform any operation of the associated instruction. The computer device 1002 also includes one or more drive mechanisms 1008 for interacting with any storage, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.

[0161] The computer device 1002 may also include an input / output module 1010 (I / O) for receiving various inputs (via input device 1012) and for providing various outputs (via output device 1014). A specific output mechanism may include a presentation device 1016 and an associated graphical user interface (GUI) 1018. In other embodiments, the input / output module 1010 (I / O), input device 1012, and output device 1014 may not be included, and the computer device 1002 may be used as a computer device in a network. The computer device 1002 may also include one or more network interfaces 1020 for exchanging data with other devices via one or more communication links 1022. One or more communication buses 1024 couple the components described above together.

[0162] The communication link 1022 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 1022 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.

[0163] Corresponds to Figures 1 to 5 The method in the embodiment of the present specification also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are executed.

[0164] The embodiment of the present specification also provides a computer-readable instruction, wherein when the processor executes the instruction, the program therein causes the processor to execute the following Figures 1 to 5 The method shown.

[0165] It should be understood that in the various embodiments of this specification, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.

[0166] It should also be understood that in the embodiments of this specification, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this specification generally indicates that the associated objects before and after are in an "or" relationship.

[0167] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this specification.

[0168] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0169] In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.

[0170] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of this specification.

[0171] In addition, each functional unit in each embodiment of this specification may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0172] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of this specification. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0173] Specific embodiments are used in this specification to illustrate the principles and implementation methods of this specification. The description of the above embodiments is only used to help understand the methods and core ideas of this specification. At the same time, for those skilled in the art, according to the ideas of this specification, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this specification.

Claims

1. A UI stylization method based on augmented reality, characterized in that: The method comprises: Obtain the augmented reality scene and UI elements to be stylized in the user's field of view; Generate, according to the augmented reality scene and the UI element to be stylized, a stylized description instruction for the UI element to be stylized, wherein the stylized description instruction is used to integrate the UI element to be stylized into the visual expression of the augmented reality scene; According to the stylized description instruction, determining each single stylized target UI element in the stylized description instruction; Rearrange the individual stylized target UI elements to generate a stylized UI component.

2. The method according to claim 1, characterized in that Generating a stylized description instruction for the UI element to be stylized according to the augmented reality scene and the UI element to be stylized includes: Inputting the visual expression corresponding to the augmented reality scene and the UI element to be stylized into the large language model to obtain the stylized description instruction output by the large language model; The large language model is obtained through pre-training.

3. The method according to claim 1, characterized in that Determining, according to the stylized description instruction, each single stylized target UI element in the stylized description instruction includes: Inputting the stylized description instruction into the text image model to obtain a single stylized target UI element output by the text image model; The text image model is obtained by training an initial text image model through stylized sample description instructions.

4. The method according to claim 3, characterized in that According to the stylized description instruction, determining each single stylized target UI element in the stylized description instruction further comprises: Convert the stylized description instructions into corresponding codes; The corresponding UI element in the stylized description instruction is determined according to the code to obtain a stylized single target UI element.

5. The method according to claim 3, characterized in that: The text image model is trained in the following way: Acquire a training sample set, the training sample set comprising: sample text and label UI elements corresponding to the sample text; Inputting the sample text into an initial text image model to obtain a single predicted UI element output by the initial text image model; A loss function is constructed according to the label UI element and the predicted UI element, and the parameters of the initial text image model are iteratively updated based on the loss function until the loss function converges to a preset threshold, thereby constructing a text image model.

6. The method according to claim 5, characterized in that After the text image model is constructed and the text image model outputs a single stylized target UI element, the method further includes: Determine a block image and text loss function according to the image features of each target UI element and the text features corresponding to the target UI element; Determining a visual feature extraction loss function according to visual features of a target image and visual features of a scene image in the visual expression, wherein the target image is obtained by combining all target UI elements; Determine the scene consistency loss function according to the features of the target image and the scene description of the scene image; At least one of the large language model and the text image model is guided to be fine-tuned according to at least one of the block image and text loss function, the visual feature extraction loss function and the scene consistency loss function.

7. The method according to claim 2, characterized in that: The large language model further outputs an arrangement layout for the augmented reality scene and the UI elements to be stylized, including: according to the augmented reality scene and the arrangement layout of the UI elements to be stylized, each single stylized target UI element is rearranged to generate a stylized UI component.

8. The method according to claim 1, characterized in that Rearrange each single target UI element to generate a stylized UI component including: Rearrange each single target UI element according to the preset arrangement layout to generate a stylized UI component.

9. The method according to claim 6, characterized in that The method further comprises: The total loss function is constructed using the following formula: + + ; in, , , They represent the weights of the block image and text loss function, the scene consistency loss function, and the normalized visual feature extraction loss function respectively; represents the block image and text loss function, represents the scene consistency loss function, Represents the normalized visual feature extraction loss function.

10. A UI stylization device based on augmented reality, characterized in that: The device comprises: An acquisition unit, used to acquire an augmented reality scene and UI elements to be stylized in the user's field of view; A first generating unit, configured to generate, according to the augmented reality scene and the UI element to be stylized, a stylized description instruction for the UI element to be stylized, wherein the stylized description instruction is used to integrate the UI element to be stylized into the visual expression of the augmented reality scene; A determination unit, configured to determine, according to the stylized description instruction, each single stylized target UI element in the stylized description instruction; The second generating unit is used to rearrange the individual stylized target UI elements to generate a stylized UI component.

11. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Application form control method, device and equipment in augmented reality space

    CN117093070A

  • Interface generation method and device, computer readable storage medium and electronic equipment

    CN118312176A

  • Browser webpage spatialization method and device

    CN119337014A

  • User Interface Creation from Screenshots

    US20180349730A1

  • Automatic User Interface Architecture

    US20200133692A1