Virtual space generation method and device, electronic equipment, medium and intelligent agent
By calling the set information to generate a large model, and generating the virtual space set information based on the item information and virtual image information, the problem of high repetition and customization costs in virtual space display is solved, and personalized and diversified virtual space display is realized, which improves the user experience.
Patent Information
- Application Number
- CN202510497487.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
In the existing virtual space display, the virtual space setting where the item is located is easily repeated, resulting in poor visual experience, and the generation cost of custom live broadcast rooms is high and the skill requirements are high.
By calling the set information to generate a large model, virtual space set information is generated based on item information and virtual image information, ensuring that the number of components and display content match the item information, and a personalized and diverse virtual space is generated.
It improves the degree of personalization and diversity of virtual space, simplifies the generation process, and improves the user experience.
Smart Images

Figure CN120374864A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as AIGC (Artificial Intelligence Generated Content). Specifically, it relates to a method, device, electronic device, medium, and intelligent agent for generating a virtual space. Background Art
[0002] With the continuous development of computer technology, it is possible to display and sell items through the Internet. For example, items can be displayed within a virtual space.
[0003] When displaying items, the setting of the virtual space where the items are located affects the viewing experience. For example, when using a single or easily repetitive virtual space to display items, the visual experience is not good. Summary of the Invention
[0004] The present disclosure provides a method, device, electronic device, medium, and intelligent agent for generating a virtual space.
[0005] According to one aspect of the present disclosure, there is provided a method for generating a virtual space, including: in response to an interaction operation, obtaining item information of an item to be displayed and virtual image information for displaying the item to be displayed; invoking a large model for generating setting information, and generating virtual space setting information according to the item information and the virtual image information, wherein the virtual space setting information includes at least one component, and the number and display content of the components match the item information; and generating a virtual space for displaying the item to be displayed according to the virtual space setting information and the virtual image corresponding to the virtual image information.
[0006] According to another aspect of the present disclosure, there is provided a device for generating a virtual space, including: an obtaining module, configured to obtain item information of an item to be displayed and virtual image information for displaying the item to be displayed in response to an interaction operation; a first generating module, configured to invoke a large model for generating setting information, and generate virtual space setting information according to the item information and the virtual image information, wherein the virtual space setting information includes at least one component, and the number and display content of the components match the item information; and a second generating module, configured to generate a virtual space for displaying the item to be displayed according to the virtual space setting information and the virtual image corresponding to the virtual image information.
[0007] According to another aspect of the present disclosure, there is provided an intelligent agent, including: an input module, configured to receive input information; a processing module, configured to determine a target task according to the input information received by the input module, determine a large model according to the target task, and execute the above information interaction method according to the large model by invoking the large model to obtain output information; and an output module, configured to output the output information obtained by the processing module.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0010] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the method as described above when executed by a processor.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 Schematically shows an exemplary system architecture that can be applied to a virtual space generation method and apparatus according to an embodiment of the present disclosure;
[0014] Figure 2 Schematically shows a flowchart of a virtual space generation method according to an embodiment of the present disclosure;
[0015] Figure 3 Schematically shows a scene schematic diagram of a virtual space according to an embodiment of the present disclosure;
[0016] Figure 4A Schematically shows an application scenario diagram for generating virtual space scenery information according to an embodiment of the present disclosure;
[0017] Figure 4B Schematically shows an application scenario diagram for generating virtual space scenery information according to another embodiment of the present disclosure;
[0018] Figure 4C Schematically shows an application scenario diagram for generating virtual space scenery information according to still another embodiment of the present disclosure;
[0019] Figure 5 Schematically shows a flowchart of a virtual space generation method according to another embodiment of the present disclosure;
[0020] Figure 6A schematic diagram of a scenario for generating new virtual space setting information according to an embodiment of the present disclosure is shown;
[0021] Figure 7 A schematic diagram of a scenario for virtual space setting information of multiple items to be displayed according to an embodiment of the present disclosure is shown;
[0022] Figure 8 A block diagram of a virtual space generation device according to an embodiment of the present disclosure is shown schematically;
[0023] Figure 9 A block diagram of the structure of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure is shown schematically; and
[0024] Figure 10 A block diagram of an electronic device suitable for implementing a virtual space generation method according to an embodiment of the present disclosure is shown schematically. Detailed implementation manners
[0025] The exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0026] In an item display scenario, it is possible to use the materials provided by an item display platform for virtual space setting. For example, when using the virtual avatar live broadcast function to display items, the materials provided by the item display platform can be used for the setting and construction of the live broadcast room. However, the number of materials provided by the item display platform is limited. For a large number of platform users, there are likely to be a large number of similar or identical live broadcast rooms, and the homogenization of live broadcast rooms is serious, lacking personalization and diversification, and the viewing experience is relatively single. In addition, users can select materials by themselves and construct live broadcast rooms without using the materials of the item display platform, but this method of generating live broadcast rooms has a high cost and requires users to have a high level of skills.
[0027] Accordingly, a first aspect of the present disclosure provides a virtual space generation method, which, in response to an interaction operation, obtains the item information of the item to be displayed and the virtual image information for displaying the item to be displayed; calls a large model for generating scene information, and generates virtual space scene information according to the item information and the virtual image information, wherein the virtual space scene information includes at least one component, and the number and display content of the components match the item information; and generates a virtual space for displaying the item to be displayed according to the virtual space scene information and the virtual image corresponding to the virtual image information. By calling the large model for generating scene information, the first aspect of the present disclosure generates virtual space scene information at the component granularity, so that the virtual space generated according to the virtual space scene information can match each item to be displayed and the virtual image, thereby enriching the personalization and diversification of the virtual space.
[0028] Figure 1 Schematically shows an exemplary system architecture to which the virtual space generation method and apparatus according to an embodiment of the present disclosure can be applied.
[0029] It should be noted that Figure 1 The example shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the virtual space generation method and apparatus can be applied may include a terminal device, but the terminal device may implement the virtual space generation method and apparatus provided by the embodiments of the present disclosure without interacting with the server.
[0030] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0031] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0032] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0033] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports the content browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0034] The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0035] It should be noted that the virtual space generation method provided by the embodiments of the present disclosure can generally be executed by terminal devices 101, 102, or 103. Correspondingly, the virtual space generation device provided by the embodiments of the present disclosure can also be set in terminal devices 101, 102, or 103.
[0036] Alternatively, the virtual space generation method provided by the embodiments of the present disclosure can generally also be executed by server 105. Correspondingly, the virtual space generation device provided by the embodiments of the present disclosure can generally be set in server 105. The virtual space generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the virtual space generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Alternatively, the virtual space generation method provided by the embodiments of the present disclosure can generally also be executed by terminal devices 101, 102, or 103. Correspondingly, the virtual space generation device provided by the embodiments of the present disclosure can generally be set in terminal devices 101, 102, or 103.
[0037] For example, the user can input the item information of the item to be displayed and the virtual image information for displaying the item to be displayed by interacting with the terminal devices 101, 102, and 103. The terminal devices 101, 102, and 103 send the obtained item information and virtual image information to the server 105. The server 105 obtains the item information and virtual image information, and calls the large model for generating scene information to generate virtual space scene information according to the item information and virtual image information, where the virtual space scene information includes at least one component, and the number and display content of the components match the item information; and generates a virtual space for displaying the item to be displayed according to the virtual space scene information and the virtual image corresponding to the virtual image information. The terminal devices 101, 102, and 103 can receive and display the generated virtual space from the server 105. Or a server or server cluster capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105 analyzes the item information and virtual image information to generate a virtual space. Or, the large model for generating scene information is set in the terminal devices 101, 102, and 103, and the terminal devices 101, 102, and 103 can directly call the local large model for generating scene information to generate virtual space scene information according to the item information and virtual image information, and generate a virtual space according to the virtual space information and virtual image information.
[0038] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0039] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc., of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0040] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.
[0041] Figure 2 Schematically shows a flowchart of a virtual space generation method according to an embodiment of the present disclosure. As Figure 2 shown, Embodiment 200 includes operations S210 to S230.
[0042] In operation S210, in response to an interaction operation, obtain the item information of the item to be displayed and the virtual image information for displaying the item to be displayed.
[0043] The item to be displayed can be an item to be displayed in a live broadcast scenario or an item to be displayed in a video scenario. The item to be displayed can be an item of multiple item types, such as clothing, daily necessities, books, electronic products, fresh fruits, etc.
[0044] For the item to be displayed, the item to be displayed can be explained through a virtual image. The virtual image information represents information related to the virtual image, such as the style information, posture information, appearance information, etc. of the virtual image. The virtual image can be a virtualized character or a video silhouette of a real person.
[0045] In one embodiment, the item information of the item to be displayed can be obtained by means of an item link, and the virtual image information can be obtained through a reference image of the virtual image.
[0046] The interaction operation can be an information acquisition operation with the terminal device. Thus, when the interaction operation is generated, the terminal device can obtain the item information of the item to be displayed and the virtual image information for displaying the item to be displayed in response to the interaction operation. For example, by interacting with the controls on the interaction page of the terminal device, the above item information and virtual image information are imported.
[0047] In operation S220, a large model for generating scene information is called, and virtual space scene information is generated according to the item information and the virtual image information.
[0048] The large model for generating scene information can be a large language model (LLMs) for processing text information or a large multimodal model (LMMs) for processing multimodal information. In this embodiment, the large model for generating scene information can be determined according to the modalities of the item information and the virtual image information.
[0049] The virtual space scene information generated by the large model for generating scene information is used to represent the scene information of the virtual space to be generated. The virtual space scene information includes the background information, color information, virtual space structure, decorations in the virtual space, etc. of the virtual space. For example, when generating a virtual space in a business style, the virtual space scene information can include decorations such as desks and bookshelves, and the color information includes black, white, gray, etc.
[0050] The virtual space scene information includes at least one component, and the number and display content of the components match the item information.
[0051] A component can be understood as a component with a display function, or a component with both display and interaction functions. For example, taking a virtual space as a live broadcast room, a component can be used to display the background information of the live broadcast room, and the decorations in the background information are regarded as elements within the background information; a component can also be a component for displaying the item title of the item to be displayed; a component can also be a component that can display the coupon information of an item and can display the coupon details through interaction; a component can also be a component that can trigger an animation effect through interaction and display the animation effect, and so on.
[0052] For multiple items to be displayed, the multiple virtual space scene information generated by calling the scene information generation large model can include the same or different numbers of components. For two virtual space scene information including the same component, the display information of the same component can be the same or different, which is related to the item information input by the item to be displayed.
[0053] For example, for component 1 used to display the item title, in the virtual space scene information of item A to be displayed, component 1 displays "Item Title A", and in the virtual space scene information of item B to be displayed, component 1 displays "Item Title B", and "Item Title A" and "Item Title B" are related to the input item information.
[0054] In operation S230, according to the virtual space scene information and the virtual avatar corresponding to the virtual avatar information, a virtual space for displaying the item to be displayed is generated.
[0055] The virtual space scene information can include the item information of the item to be displayed and / or the item to be displayed. After generating the virtual space scene information, the virtual avatar can be integrated with the virtual space scene information to obtain a virtual space. In this virtual space, the virtual avatar can display the item to be displayed by means of explanation.
[0056] For example, the virtual avatar can be placed in a preset position area so that in the virtual space generated according to the virtual space scene information and the virtual avatar, there is no occlusion or as little occlusion as possible between the virtual avatar and the components in the virtual space scene information.
[0057] In the embodiments of the present disclosure, since the virtual space scene information is generated based on the item information and the virtual avatar information, the virtual space scene information can be adapted to each item to be displayed and the virtual avatar, making the scene of the virtual space more personalized and diversified, and providing a good user experience. For a single item to be displayed, the virtual space scene information generated by the scene information generation large model includes at least one component, and the number and display content of the components match the item information. Thus, the virtual space scene information can more accurately display the item to be displayed at the component granularity, improving the adaptation degree between the virtual space and the item to be displayed, and further enriching the personalization and diversification of the virtual space. In addition, through simple interaction operations and by invoking the scene information generation large model to generate the virtual space scene information, the generation process of the entire virtual space is simple and more intelligent, providing a good user experience.
[0058] Figure 3 Schematically shows a scene diagram of a virtual space according to an embodiment of the present disclosure. As Figure 3 shown, in the virtual space 310 of Embodiment 300, the virtual space scene information may include background information 313, and the ornaments 311 and 312 included in the background information 313. The ornament 312 is used to carry the virtual avatar 330, such as steps. In the virtual space 310, there are also multiple virtualized components 320. For example, the component 320 includes a component for displaying "Item Information 1", and three components for displaying "Item Function Point 1" to "Item Function Point 3". The item information 1 may be a link, price, schematic diagram, and item title of item A.
[0059] In another embodiment, for the item to be displayed "Electronic Product XX", a new component may also be included on the right side of the virtual space to display the warranty information of Electronic Product XX, such as displaying "Free repair and replacement within 1 year".
[0060] According to an embodiment of the present disclosure, the virtual avatar information includes the pose information of the virtual avatar; by invoking the scene information generation large model, according to the item information and the virtual avatar information, virtual space scene information is generated, including: invoking the scene information generation large model, according to the item information and the pose information, generating at least one initial component, where the number of the initial components matches the item information; updating the initial components according to the item information to obtain components; and generating the virtual space scene information according to at least one component.
[0061] The pose information is used to characterize the action of the virtual avatar. For example, the pose information may include a standing pose, a sitting pose, a squatting pose, etc. In this embodiment, the item information and the pose information can be input into the scene information generation large model.
[0062] Considering the virtual avatar display scenario, the virtual space also includes a virtual avatar. If the difference between the virtual avatar and the virtual space scenery information in the virtual space is too large, the visual experience is not good. Therefore, when generating the virtual background scenery information, the actions of the virtual avatar are also used as the input to the large model for generating scenery information, so that the large model for generating scenery information can generate virtual space scenery information that is consistent with the style of the virtual avatar and the item to be displayed, improving the visual experience.
[0063] The large model for generating scenery information is used to perform component generation tasks and component update tasks. The component generation task can be understood as a matching task. The large model for generating scenery information can match at least one initial component from the component library according to the item information and pose information. Then, through the component update task, the matched initial component is updated according to the item information to obtain a component that matches the item information. The number of initial components can be more than, less than, or the same as the number of components in the virtual space scenery information.
[0064] For example, when performing the component generation task, the large model for generating scenery information can analyze the item information and pose information, determine the component category that matches the pose information and item information, and then determine the initial component from the component library according to the component category. In one embodiment, there can be multiple initial components under the same component category in the component library, and the initial component can be randomly determined from the component library according to the component category. By random matching, on the basis of ensuring the personalization of matching the item information, the diversity of the generated components is further ensured. When performing the component update task, the corresponding display content can be extracted from the item information, and then the display content is filled into the initial component, or the initial display content in the initial component is replaced with the display content to obtain the component.
[0065] For example, the initial component carries position information, and the component obtained after updating the display content still carries this position information. Thus, after obtaining at least one component, the at least one component can be assembled according to the position information carried by each component to obtain the virtual space background information.
[0066] For another example, take the posture information as sitting position and the item to be displayed as a down jacket of brand A. After inputting the item information and the posture information into the scene information generation large model, the scene information generation large model obtains from the item information including "price", "winter", and "lowest price of the day". When the execution component of the scene information generation large model generates a task, it can match the initial components according to the item information and "sitting position": the first component for displaying the background information, the second component for displaying the item title, and the third component for displaying the item link of the item to be displayed. The background information displayed by the first component is a cloakroom, and the clothes in the cloakroom are winter-related clothes. In addition, the first component also includes a table, which matches the "sitting position". The item title displayed by the second component is "Special winter purchase, the down jacket of brand A has the lowest price today". The third component displays the item link, and the third component includes the price and image of the item to be displayed.
[0067] In an embodiment of the present disclosure, the scene information generation large model generates at least one initial component according to the item information and the posture information, which can ensure that the generated initial components are adapted to both the posture information and the item information at the same time, and avoid a large difference between the virtual image and the virtual space scene information. After determining the initial components, the initial components are updated according to the item information to obtain components; according to at least one component, the virtual space scene information is generated, which can further ensure that each component in the virtual space scene information can match the item information from the component granularity, avoid scene deviation in the details, and thus improve the matching degree between the virtual space scene information obtained according to at least one component and the item information.
[0068] Figure 4A Schematically shows an application scenario diagram of generating virtual space scene information according to an embodiment of the present disclosure. As Figure 4A shown, in embodiment 400A, the item information 401 of the item to be displayed and the posture information 402 of the virtual image are input into the scene information generation large model M1. The scene information generation large model M1 executes a component generation task and generates at least one initial component according to the item information 401 and the posture information 402, such as initial components 403-1...; executes a component update task and updates each initial component according to the item information 401 to obtain at least one component, such as components 404-1.... For example, the initial component 403-1 is updated according to the item information 401 to obtain the component 404-1. According to at least one component, such as components 404-1..., the virtual space scene information 405 is generated and output.
[0069] In this embodiment, a first prompt message can be generated based on the item information 401 and the pose information 402, and the first prompt message is input into the scene information generation large model M1, so as to obtain the virtual space scene information 405. The first prompt message can be: Step 1: Please analyze the provided item link and the pose of the virtual object below and match the appropriate components; Step 2: Please fill in the content of the components matched in the first step according to the following item information to obtain the updated components. Please output the combined information of the components obtained in the second step.
[0070] According to an embodiment of the present disclosure, at least one initial component matching the item information is generated according to the item information and the pose information, including: estimating the pose of the virtual image according to the pose information to obtain the virtual image position information; and generating at least one initial component matching the item information according to the item information, the virtual image position information and the pose information.
[0071] In one embodiment, the pose information of the virtual image can be the pose information within a preset time period, such as the pose information of the virtual image within 10s; or, the pose information can also be the pose information of the virtual image under a plurality of predetermined actions.
[0072] The virtual image position information represents the position change range of the virtual image in the virtual space. For example, in the virtual space, the virtual image can perform actions such as "shaking" and "raising hands", and the positions of various actions all fall within this position change range.
[0073] Since in the virtual space, the pose of the virtual image is usually changing. For example, when in a standing or sitting posture, the virtual image can display items through hand movements. Therefore, in order to ensure that the virtual image does not block or blocks as few components as possible when changing its pose, when the scene information generation large model generates the initial components, it is necessary to take into account the position relationship between the initial components and the virtual image. In addition, when generating the initial components, the scene information generation large model can also take into account the position relationship between the initial components.
[0074] In one embodiment, the pose of the virtual image can be estimated based on methods such as template matching, deep learning, and reinforcement learning according to the pose information to obtain the virtual image position information. For example, when the virtual image is a real person, the pose can also be estimated based on a bone model. For example, by analyzing the actions of the virtual image and calculating the angles and positions of the joints according to the kinematic principles of the bone model, the pose is estimated to obtain the virtual image position information.
[0075] In one embodiment, generating at least one initial component that matches the item information based on the item information, virtual avatar position information, and pose information can be achieved in various ways. For example, when training the large model for generating set design information, the positional relationships between components and the positional relationships between the virtual avatar and the components (including layout, perspective, and depth information) are annotated in the training data. Through training, the large model can generate at least one initial component with non-overlapping or minimally overlapping components.
[0076] Alternatively, appropriate prompt information is constructed based on the item information, virtual avatar position information, and pose information. This prompt information is used to indicate the generation of at least one initial component with non-overlapping or minimally overlapping components. Under the guidance of this prompt information, the large model for generating set design information is utilized to generate at least one initial component. For example, the prompt information can be specific layout rules or constraints, such as no positional coincidence between the virtual avatar and the initial component and / or the proportion of the overlapping area being less than a certain predetermined proportion. For instance, layout optimization algorithms (such as force-directed method, position constraint optimization) can be employed to generate the layout rules or constraints.
[0077] In the embodiments of the present disclosure, by performing pose estimation on the virtual avatar according to the pose information, the virtual avatar position information is obtained; and at least one initial component that matches the item information is generated based on the item information, virtual avatar position information, and pose information, such that there is no occlusion or minimal occlusion between the generated initial components and / or between the initial components and the virtual avatar, thereby ensuring a reasonable arrangement between the components determined based on the initial components and between the components and the virtual avatar, and improving the aesthetic degree of the virtual space and the usage experience. In a specific embodiment, generating at least one initial component based on the item information, virtual avatar position information, and pose information includes: generating at least one initial component that matches the item information based on the item information and pose information in an area different from the virtual avatar position information; and when the positions between at least one initial component do not meet the predetermined position conditions, adjusting the positions between at least one initial component until at least one initial component that meets the predetermined positional relationship is generated.
[0078] In this embodiment, since the virtual avatar position information of the virtual avatar has been determined through pose estimation, in order to reduce the complexity of the generation process, the area different from the virtual avatar position information can be directly used as a position limit, and in the area different from the virtual avatar position information, at least one initial component that matches the item information is generated based on the item information and pose information.
[0079] Regarding the positional relationship between the initial components, the scene information generation large model can normally generate the initial components, and the generated initial components carry position information (the output of position information can be requested in the prompt information). After generating the initial components, the physical engine can be used to simulate the positions between the initial components and the virtual avatar, and by calling the collision detection tool, according to the position information of the initial components, it can be determined whether the positions of at least one initial component meet the predetermined position conditions. In the case where the positions do not meet the predetermined position conditions, according to the respective position information of the initial components, the positions of one or more initial components are adjusted until at least one initial component that meets the predetermined position conditions is obtained. Alternatively, the predetermined position conditions can be used as prompt information, and the scene information generation large model is required to regenerate at least one initial component that meets the predetermined position conditions according to the predetermined position conditions, item information, and pose information. In this embodiment, the predetermined position conditions can be that there is no overlap in the position information of the initial components, or the proportion of the overlapping area is less than a certain predetermined proportion.
[0080] In one embodiment, the positions of the initial components can also be adjusted according to the spatial arrangement relationship corresponding to the predetermined viewing angle, so that the initial components not only meet the predetermined positional relationship but also meet the spatial arrangement relationship. For example, the predetermined viewing angle can be the viewing angle of viewing the virtual space at the central position of the virtual space. Or, the predetermined viewing angle can also be the viewing angle determined based on historical viewing habits. The spatial arrangement relationship of the predetermined viewing angle can be: when viewing the virtual space from the predetermined viewing angle, the spatial arrangement relationship between the components and / or between the components and the virtual avatar.
[0081] For example, when the predetermined viewing angle is to view the virtual space at the central position of the virtual space, the spatial arrangement relationship can be that the components in the virtual space are symmetric left and right and / or up and down. By adjusting the positions of the initial components according to the spatial arrangement relationship corresponding to the predetermined viewing angle, the adjusted virtual space can further meet the visual experience.
[0082] In the embodiments of the present disclosure, by generating at least one initial component that matches the item information according to the item information and pose information in an area different from the position information of the virtual avatar; in the case where the positions between at least one initial component do not meet the predetermined position conditions, adjusting the positions between at least one initial component until at least one initial component that meets the predetermined positional relationship is generated, it is possible to generate initial components that do not overlap or overlap as little as possible between the initial components and between the initial components and the virtual avatar while reducing the complexity of the generation process, and improve the aesthetic degree of the virtual space and the usage experience.
[0083] In another embodiment, after constructing a virtual space based on the generated virtual space scene information, it is possible to detect whether there is occlusion by calling a collision detection tool, and adjust the position when there is occlusion. The adjustment can be performed in the above-mentioned manner.
[0084] According to an embodiment of the present disclosure, at least one initial component that matches the item information is generated based on the item information and the pose information, including: extracting item sub-information of at least one display dimension and the item type of the item to be displayed from the item information; and generating an initial component that matches the item sub-information of each display dimension according to the pose information and the item type.
[0085] The item information of the item to be displayed may include at least one item sub-information, and at least one item sub-information corresponds to one display dimension. The display dimension can be understood as the dimension that needs to be displayed in the virtual space. For example, the display dimension may include at least one of the following: item title, item function, coupon, logistics guarantee, item after-sales, etc. The item function may be the highlight function of the item to be displayed, etc. In one embodiment, for the display dimension of the item function, it may include multiple item sub-informations, such as the warmth retention performance and fabric of a down jacket.
[0086] The item type is used to characterize the type to which the item to be displayed belongs, such as: clothing, daily necessities, books, electronic products, fresh fruits, etc.
[0087] The scene information generation large model can also perform a content analysis task to perform content analysis on the item information, so as to extract at least one item sub-information of the display dimension and the item type. After that, the scene information generation large model can also perform a component generation task to generate an initial component that matches each item sub-information under each display dimension according to the item type, the pose information, and at least one item sub-information of each display dimension. The task description of the content analysis task can be constructed in the prompt information of the scene information generation large model to extract at least one item sub-information of the display dimension and the item type.
[0088] For example, the component category can be determined according to the item type, the pose information, and the label of the display dimension, and then the initial component can be determined from the component library according to the component category and the item sub-information. Or, directly determine the component category according to the item type, the pose information, the display dimension, and the label of the item sub-information, and then determine the initial component from the component library according to the component category. In one embodiment, the initial component can be randomly determined from the component library according to the component category, and the diversity of the generated components can be further ensured on the basis of ensuring personalization through random matching.
[0089] In an embodiment of the present disclosure, a large model for generating scene information performs content analysis on item information to generate item sub-information and item types of at least one display dimension, enabling users to directly input item information according to interaction operations without manually sorting and analyzing item sub-information, which is simple and convenient to use. In addition, when generating initial components, by combining pose information, display dimensions, and item types, initial components with a unified style can be generated more precisely, making the generation process of the virtual space scene information more controllable, and the detailed parts of the component granularity can match the item sub-information more precisely.
[0090] Figure 4B FIG. schematically shows an application scenario diagram for generating virtual space scene information according to another embodiment of the present disclosure. In Embodiment 400B, the operations of item information 401, pose information 402, and generating at least one component 404-1... are similar to those in Embodiment 400A and will not be described herein again.
[0091] In this embodiment, the large model M1 for generating scene information can perform a content analysis task to determine the item type 406 and at least one item sub-information, such as item sub-information 407-1..., based on the item information 401. The at least one item sub-information can belong to at least one display dimension. Subsequently, the large model M1 for generating scene information can perform a component generation task to determine the initial component 403-1 corresponding to the item sub-information 407-1 based on the item type 406, the item sub-information 407-1, the display dimension of the item sub-information 407-1, and the pose information 402. The operations for performing the component update task are similar to those in Embodiment 400A and will not be described herein again.
[0092] In this embodiment, second prompt information can be generated based on the item information 401 and the pose information 402, and the second prompt information is input into the large model M1 for generating scene information to obtain the virtual space scene information 405. The second prompt information can be: Step 1: Please analyze the provided item link below, extract the item type and item sub-information of at least one display dimension from the item link, and match a suitable component according to the information extracted above and the pose of the virtual object; Step 2: Please perform content filling on the component matched in the first step according to the following item information to obtain an updated component. Please output the combined information of the components obtained in the second step.
[0093] According to an embodiment of the present disclosure, generating initial components that match the sub-information of an item for each display dimension based on pose information and item type includes: when the component library associated with the scene information generation large model does not include the item type, generating initial components that match the sub-information of an item for each display dimension based on the pose information and candidate item types similar to the item type in the component library; when the component library does not include the sub-information of the item for the display dimension, determining candidate sub-information of the item for at least one candidate display dimension according to the evaluation information included in the item information; and generating initial components that match the candidate sub-information of the item for each candidate display dimension based on the pose information and the item type.
[0094] Considering that the item to be displayed may have new display dimensions, or the item type of the item to be displayed belongs to an item type of information that the scene information generation large model may not have learned such features, and the component library associated with the scene information generation large model may also not include the display dimensions and / or new item types of the above information. In this case, by determining whether the component library includes the above display dimensions and / or item types, it is determined whether the extracted display dimensions and / or item types are new information, and when it belongs to new information, the initial components of the item to be displayed are generated using the existing components in the component library.
[0095] The scene information generation large model can be associated with a component library, which includes not only multiple components of items belonging to multiple item types, for example, multiple components of items in the clothing category, beauty category, and tourism category; but also can include multiple components for each display dimension, for example, multiple components for display dimensions such as coupons, logistics guarantee, and item after-sales. In one embodiment, the components in the component library can be pre-determined or newly generated by the scene information generation large model.
[0096] In one embodiment, the first similarity between the item type and the item types in the component library, and / or the second similarity between the item type and the type description information of the item types in the component library can be calculated, and the item types with the first similarity and / or the second similarity exceeding the corresponding similarity thresholds are used as candidate item types. When the candidate item types are determined, initial components that match the sub-information of the item for each display dimension can be generated based on the pose information and the candidate item types, and the generation operation is similar to the above and will not be elaborated here.
[0097] In another embodiment, for display dimensions not included in the component library, the large language model can be used to perform content analysis on the evaluation information contained in the item information to extract the first dimension description information of the display dimension. By calculating the third similarity between the first dimension description information and the display dimensions in the component library and / or the fourth similarity between the first dimension description information and the second dimension description information of the display dimensions in the component library, and taking the display dimensions with the third similarity and / or the fourth similarity exceeding the corresponding similarity threshold as candidate display dimensions, and taking the extracted item sub-information as the candidate item sub-information under the candidate display dimensions. After that, the initial components matching the candidate item sub-information of each candidate display dimension can be generated according to the pose information and the item type. The generation operation is similar to the above and will not be elaborated here.
[0098] The evaluation information can be feedback information on after-sales service, comment information on public platforms, market analysis of items, etc.
[0099] In addition, for display dimensions and / or item types not included, the large model for generating scene information can also be fine-tuned using the information of the new display dimensions and / or item types, so that the generating large model learns similar display dimensions and / or item types from the existing component library, adaptively generates the initial components of the display dimension and / or item type, and stores the generated initial components in the component library. Thus, the initial components of the display dimension and / or item type are continuously enriched in the component library to improve the subsequent generation efficiency.
[0100] For example, the fine-tuning process can be carried out in the following way: for the new display dimension, collect the comment information of this display dimension. For the new item type, collect item pictures, item descriptions, comment information, etc. under this item category, analyze the industry trends and user preferences of this item type, and / or analyze the currently popular scene color systems and styles to guide the generation of new components under this item category. Mark the key information of the new display dimension and item type, and extract the features related to scene design, such as color tendency, style features, key selling points, etc. (this feature corresponds to the content extraction, image generation, scene color system generation, etc. that the above-mentioned large model for generating scene information can achieve). For example, take the information related to the extracted features as the key information. In addition, as a supplement, natural language processing technology can also be used to analyze the information in the above text modality (such as comment information, item description, etc.), extract the semantic information useful for scene design, and synchronously use this speech information as the key information mentioned above to extract the features related to scene design.
[0101] For the newly extracted features, the large model adaptively learns similar color systems, styles, and other features from the existing component library, and adjusts and optimizes according to the new features to generate new initial components. The new initial components can be dynamically added to the component library. In this process, transfer learning technology can be used to transfer the knowledge of the existing model to the new item types and display dimensions to accelerate the fine-tuning speed.
[0102] In addition, for the new display dimensions and item types, it is also possible to collect user feedback on the new display dimensions and item types, and continuously optimize the component library and the large model. During the process of collecting feedback information, for the scene information generation large model that can generate new display dimensions and item types, small-scale experiments and A / B testing are carried out online to obtain feedback on the components and scene color systems to ensure their effectiveness and visual effects in actual applications.
[0103] It should also be noted that in the case where the component library does not include item types and display dimensions, the above two methods can be combined to generate initial components, which will not be elaborated here.
[0104] In the embodiments of the present disclosure, the information of the candidate item types and candidate display dimensions in the component library associated with the scene information generation large model is used, so that the scene information generation large model can generate initial components for the new display dimensions and / or new item types, meeting diverse generation requirements and improving the user experience.
[0105] According to the embodiments of the present disclosure, the initial components further include the background image of the virtual space.
[0106] In one embodiment, the scene information generation large model can perform the text-to-image task. For example, the scene information generation large model can generate an image description that matches the pose information and item information, and then the background information large model can generate a background image according to the image description, and the background image can be used as the display background in the virtual space of the item to be displayed.
[0107] For example, the set design information generation large model can determine information such as the color, style, size, and function of an item based on the item information, and extract a first background description that matches the item characteristics; and / or, it can also determine the visual style and brand image of the brand to which the item belongs based on the item information, and extract a second background description that matches the brand characteristics. According to the pose information and the item information, a third background description for describing the display scene of the item to be displayed is generated; and based on the third background description, information such as emotion and atmosphere is extracted to generate a fourth background description for adjusting the set color system and style. In addition, as a supplement, the set design information generation large model can also generate a fifth background description that matches the user's preferences based on the user's historical operation information. For example, the historical operation information can be browsing history, etc. Or, it can also generate a sixth background description that matches the current trend based on the design or color information with a high discussion degree in the current social media. Or, it can also generate a seventh background description that matches the festival or e-commerce activity during the sales time and an eighth background description that matches the regional characteristics or text characteristics of the sales location based on information such as the sales time and sales location of the item information. By combining at least one of the above first to eighth background descriptions, an image description can be generated.
[0108] In another embodiment, the method further includes: determining at least one historical background image according to the pose information and the item type, where the historical background image includes a background image generated using the same pose information and the same item type; generating a background image that matches the item information according to the at least one historical background image and the item information.
[0109] The component library associated with the set design information generation large model may include historical background images. At least one historical background image can be matched from the component library according to the pose information and the item type of the item to be displayed. For example, when there are tags of pose information and item type in the background images in the component library, the above historical background images can be determined through the tags.
[0110] After determining at least one historical background image, the at least one historical background image can be used as a reference image, and a new background image can be generated according to the item information. In one embodiment, if there are multiple historical background images, the most recent historical background image can be selected as the reference image and a new background image can be generated according to the item information.
[0111] In the embodiments of the present disclosure, by using the historical background images generated with the same pose information and the same item type according to the pose information and the item type, and generating a background image that matches the item information according to the at least one historical background image, a more accurate background image can be generated with reference to the historical background image, enriching the user's viewing experience of the background image in the virtual space.
[0112] In one embodiment, when the initial component is a background image, the background image can be directly used as the final component, that is, the initial component is the final component.
[0113] According to an embodiment of the present disclosure, the item information includes at least one item image information; generating an initial component that matches the item sub-information of each display dimension according to the pose information and the item type includes: determining target item image information from at least one item image information; determining a background color system from multiple colors according to the proportions of multiple colors in the target item image information; and generating an initial component that matches the item sub-information of each display dimension according to the pose information, the item type, and the background color system.
[0114] The item image information is used to reflect the item information of the item to be displayed from the image perspective. For example, at least one item image information can be at least one item to be displayed from at least one display angle, such as at least one item image information displayed in the item sales link. The target item image information can be the item image information at the top of the item sales link, or the item image information marked as the main image in the item sales link.
[0115] The background color system can be a color range including multiple similar colors. The background color system is used to characterize the colors optional for each component in the virtual space background information, and the colors of at least one component can be the same or different. In an embodiment of the present disclosure, the background color system can be determined according to the item image information of the item to be displayed, so that the color of the subsequently generated initial component is consistent with the overall color of the item to be displayed.
[0116] In one embodiment, the item image information can be an item picture, and the target item image information can be a target item picture, such as the front picture of the item. The background information generation large model can determine the color to which each pixel point in the target item picture belongs. Since the pixel points belonging to the same color range can all be regarded as the same color, the proportions of the respective colors in the target item picture, that is, the proportions of multiple colors, can be determined according to the number of pixel points of each color. Then, the color with the largest number of pixel points can be used as the background color system.
[0117] For the already determined background color system, the background information generation large model can perform a component generation task, and generate an initial component that matches the item sub-information of each display dimension according to the pose information, the item type, and the background color system. Then, the background information generation large model performs a component generation task, determines the component category according to the pose information, the item type, the background color system, the display dimension, and the label of the item sub-information, and then randomly determines an initial component from the component library according to the component category.
[0118] In an embodiment of the present disclosure, since a background color system is newly added in the process of determining the initial components, and this background color system is determined according to the item image information of the item to be displayed, the color systems of multiple components in the entire virtual space are ensured to be unified, avoiding a poor visual experience caused by large color differences among the components and improving the display effect of the virtual space.
[0119] Figure 4C FIG. schematically shows an application scenario diagram for generating virtual space scene information according to another embodiment of the present disclosure. In Embodiment 400C, the item information 401, the pose information 402, and the operation of generating at least one component 404-1... are similar to those in Embodiment 400B, and the operation of performing the content analysis task is similar to that in Embodiment 400B, which will not be elaborated here.
[0120] In this embodiment, the scene information generation large model M1 can perform an image analysis task and determine the background color system 408 according to at least one item image information in the item information 401. The scene information generation large model M1 can perform a component generation task and determine the initial component 403-1 corresponding to the item sub-information 407-1 according to the background color system 408, the item type 406, the item sub-information 407-1, the display dimension of the item sub-information 407-1, and the pose information 402.
[0121] In a specific embodiment, determining the background color system according to the proportions of multiple colors in the target item image information includes: determining the background color system according to the proportions of multiple colors in the background image and the proportions of multiple colors in the target item image information.
[0122] The initial component may include a background image, and the background image also includes multiple colors. Therefore, the scene information generation large model can determine the color to which each pixel point in the background image belongs, and can determine the proportion of each color in the background image according to the number of pixel points of each color, that is, the proportions of multiple colors in the background image. Then, the color with the largest number of pixel points in the background image and the target item image is used as the background color system.
[0123] In the embodiment of the present disclosure, by comprehensively determining the background color system according to the proportions of multiple colors in the background image and the proportions of multiple colors in the target item image information, the source of the background color system can be enriched, so as to generate a virtual space with unified colors and improve the user experience.
[0124] Figure 5 FIG. schematically shows a flowchart of a virtual space generation method according to another embodiment of the present disclosure. As Figure 5 shown, Embodiment 500 includes operation S210, operations S510 to S520, operation S230. Operations S210 and S230 are the same as or similar to those in Embodiment 200, which will not be elaborated here.
[0125] In operation S510, custom information is obtained.
[0126] The custom information includes setting style information and / or a custom color system. The setting style information is information used to describe the style of the virtual space setting. For example, the setting style information can be a sports style, a business style, a leisure style, a fresh style, a natural scenery style, etc. The custom color system can be the color system of the virtual space determined through a custom operation.
[0127] For example, the custom information can be determined through a selection operation or an input operation. Or, it can also be determined according to historical usage information.
[0128] In one embodiment, operation S510 and operation S210 can be implemented simultaneously. For example, in response to an interaction operation, item information of the item to be displayed, virtual image information for displaying the item to be displayed, and custom information are obtained. For example, through the same interaction interface, item information, virtual image information, and custom information can be obtained in response to at least one interaction operation.
[0129] In operation S520, a large model for generating setting information is called, and virtual space setting information is generated according to the custom information, item information, and virtual image information.
[0130] The large model for generating setting information can generate at least one initial component according to the custom information, item information, and virtual image information; update the initial component according to the item information to obtain a component; and generate virtual space setting information according to at least one component. In this embodiment, the large model for generating setting information can execute a component generation task to generate at least one initial component according to the custom information, item information, and virtual image information; execute a component update task to update the initial component according to the item information to obtain a component. Then, virtual space setting information is generated according to at least one component. The component generation task and the component update task are similar to those above and will not be elaborated here.
[0131] In one embodiment, after the large model for generating setting information executes a content analysis task, it can execute a component generation task to generate an initial component that matches the sub-item information of the item in each display dimension according to the pose information, item type, and custom information (custom color system and / or setting style information).
[0132] In another embodiment, when the item information includes at least one item image information and the input of the large model for generating setting information includes a custom color system, the large model for generating setting information can generate an initial component according to the pose information, item type, and custom color system; or generate an initial component according to the pose information, item type, and color system of the setting.
[0133] Alternatively, initial components are generated based on pose information, item type, background color system, and custom color system. For example, a new background color system can be determined based on the background color system and the custom color system, and initial components can be generated based on the pose information, item type, and the new background color system. For instance, in the color system sequence arranged according to pixel values, if the background color system and the custom color system are adjacent, the background color system and the custom color system are juxtaposed as the new background color system; if the background color system and the custom color system are not adjacent, one of the background color system and the custom color system can be randomly determined as the new background color system.
[0134] In an embodiment of the present disclosure, after obtaining custom information, the custom information, item information, and virtual image information are all used as inputs for calling the large model for generating background information, supporting the generation of virtual space background information according to the user-defined color system and / or custom background style information, thereby generating a virtual space that better meets the user's needs, with simple operations and improved personalization and flexibility of the generated virtual space.
[0135] According to an embodiment of the present disclosure, the virtual space generation method further includes: in response to an interaction operation on the virtual space background information, obtaining custom information; calling the large model for generating background information, and updating the virtual space background information according to the custom information to obtain new virtual space background information.
[0136] In this embodiment, virtual space background information that meets the user's needs can be generated through multiple rounds of operations. For example, after the large model for generating background information generates virtual space background information based on item information and virtual image information, if there are differences between the generated virtual space background information and the user's needs, the user can generate new virtual space background information through the next round of optimization operations.
[0137] The interaction operation on the virtual space background information can be an update operation on the virtual space background information. In one embodiment, in response to the interaction operation on the virtual space background information, an update interface can be displayed, and custom information can be obtained through interaction with the update interface. The custom information can include a custom color system and / or background style information.
[0138] In one embodiment, a virtual space without a virtual image can be generated based on the virtual space background information and presented to the user, thereby presenting the actual effect after rendering the virtual space background information to the user, facilitating the user to intuitively understand whether the generated virtual space background information meets the user's needs.
[0139] In yet another embodiment, the update page and the virtual space background information generated in the previous round can also be displayed simultaneously, facilitating targeted modification with reference to the virtual space background information generated in the previous round.
[0140] It can be understood that after generating new virtual space scene information, if the user requirements are still not met, the above operation of updating the virtual space scene information can be repeated until new virtual space scene information that meets the user requirements is obtained. After determining the new virtual space scene information, the final virtual space can be generated based on the new virtual space scene information and the virtual image.
[0141] Call the large model for generating scene information to execute the scene update task, and update the virtual space scene information according to the custom information to obtain new virtual space scene information. In the scene update task, each component in the virtual space scene information can be updated according to the custom information to obtain new components; then, new virtual space scene information can be generated based on at least one new component.
[0142] In the embodiments of the present disclosure, by first generating virtual space scene information based on the item information and the virtual image information, then obtaining custom information in response to an interaction operation on the virtual space scene information, and calling the large model for generating scene information to update the virtual space scene information according to the custom information, it is possible to first generate diverse virtual space scene information and then perform targeted updates on the details according to the custom information, so as to ensure that the updated virtual space scene information will be more in line with the user requirements in terms of details and improve the usage experience.
[0143] Figure 6 A schematic diagram of a scenario for generating new virtual space scene information according to an embodiment of the present disclosure is schematically shown.
[0144] As Figure 6 shown, Embodiment 600 includes an update interface 610 and a virtual space 620 generated based on the virtual space scene information XY, where the virtual space 620 does not include a virtual image.
[0145] The update interface 610 includes an item identifier 611 of the item to be displayed, a plurality of color system controls 612 available for selection, a plurality of style controls 613 available for selection, and a generation control 614. In the update interface 610, the user can determine the custom color system and / or scene style information by interacting with the plurality of color system controls 612 and the plurality of style controls 613. The style control 613 may include an identifier of the style, such as "Style 1", a description of the style "Description 1", and an example of the style "Style illustration".
[0146] For example, the user can select Color 2 as the custom color system by clicking on Color 2, and select Style 1 as the scene style information by clicking on Style 1. Then, by clicking on the generation control 614, an interaction operation for the virtual space scene information XY is generated to obtain custom information.
[0147] According to an embodiment of the present disclosure, the virtual space generation method further includes: when obtaining the item information and virtual image information of multiple items to be displayed, calling a large model for generating scene information, and generating virtual space scene information according to the multiple item information and virtual image information; wherein, the virtual space scene information includes virtual space scene sub-information matching each item to be displayed, and each virtual space scene sub-information includes at least one component matching the item information of each item to be displayed, and the multiple virtual space scene sub-information includes at least one background information.
[0148] For a single item to be displayed, calling a large model for generating scene information can generate corresponding virtual space scene information for the single item to be displayed. However, for multiple items to be displayed that need to be displayed within the same time period, they are in the same virtual space. For example, when displaying items through a live broadcast room, multiple items can be displayed in a single live broadcast. In this scenario, if virtual space scene information is generated for each item to be displayed separately, it may cause the scenes of multiple items to be displayed in the same virtual space to be too fragmented, resulting in a poor visual viewing experience.
[0149] Therefore, for multiple items to be displayed, the respective item information and virtual image information of the multiple items to be displayed are used as the input of the large model for generating scene information, so that the large model for generating scene information analyzes the relationships between the multiple items to be displayed and generates a set of virtual space scene information for the multiple items to be displayed.
[0150] In one embodiment, for each item to be displayed, a large model for generating scene information is called, and virtual space scene sub-information for each item to be displayed is generated according to the item information and virtual image information of the item to be displayed. At least one target virtual space scene sub-information is determined from the background information of the multiple virtual space scene sub-information, and the large model for generating scene information is called to modify the background information of the multiple virtual space scene sub-information to the background information of the at least one target virtual space scene sub-information, so that the multiple virtual space scene sub-information includes at least one background information.
[0151] For example, the background information of the multiple virtual space scene sub-information can be modified to the same background information; or, one target virtual space scene sub-information is determined from the background information of the virtual space scene sub-information of each item type, and the background information of the virtual space scene sub-information of the same item type is modified to the background information of the above target virtual space scene sub-information.
[0152] In another embodiment, at least one target item to be displayed can also be determined from multiple items to be displayed, and a large model for generating scene information is called. Target virtual space scene sub-information is generated based on the item information and virtual image information of each target item to be displayed. Then, for each item to be displayed that is not a target item to be displayed, a large model for generating scene information is called, and virtual space scene sub-information with the same background information as that of a certain target virtual space scene sub-information is generated based on the item information, virtual image information of each item to be displayed, and the background information of at least one target virtual space scene sub-information. For example, the same background information is used for the same item type.
[0153] It should be noted that when calling the large model for generating scene information, multiple virtual space scene sub-information including at least one background information can be generated by modifying the prompt information of the large model for generating scene information.
[0154] In the embodiments of the present disclosure, by calling the large model for generating scene information, virtual space scene information is generated based on multiple item information and virtual image information, and multiple virtual space scene sub-information includes at least one background information, thereby avoiding the scene of multiple items to be displayed in the same virtual space from being too fragmented and improving the usage experience.
[0155] According to the embodiments of the present disclosure, calling the large model for generating scene information and generating virtual space scene information based on multiple item information and virtual image information includes: calling the large model for generating scene information, and determining the item type and / or color system of the scene of each item to be displayed from the item information of each item to be displayed; determining the target item type and / or target color system of the scene according to the number of items to be displayed belonging to the same item type and / or the same color system of the scene, and determining the target item type and / or target color system of the scene as the normalization information; and generating virtual space scene information based on the normalization information, multiple item information, and virtual image information.
[0156] For example, multiple items to be displayed can respectively belong to multiple item types such as clothing and daily necessities, and the color system of the scene of each item to be displayed can be determined according to the item image information of each item to be displayed. There can be multiple color systems for multiple items to be displayed.
[0157] The normalization information can be a certain item type and / or a certain color system of the scene determined from multiple item types. For example, the item type of item to be displayed A and / or the color system of the scene of item to be displayed B can be determined as the normalization information.
[0158] In one embodiment, the item type with the largest number of items among multiple items to be displayed can be determined as the target item type; and / or, for the item type with the largest number of items, an item to be displayed can be randomly determined from this item type, and the background color system of this item to be displayed can be used as the target background color system. The normalization information can include the above-mentioned target item type and / or target background color system.
[0159] For example, for each item to be displayed, the normalization information can be used as the prompt information for the large model for generating the set scene information. The large model for generating the set scene information is called, and according to the normalization information, the item information, and the virtual image information, the virtual space set scene sub-information is generated.
[0160] In the embodiments of the present disclosure, according to the item type and / or the background color system, the virtual space set scene sub-information of multiple items to be displayed is unified into at least one background information, so as to ensure the personalization and diversity of the virtual space set scene information while avoiding the set scenes of multiple items to be displayed in the same virtual space from being too fragmented, and improving the usage experience.
[0161] Figure 7 A schematic diagram of the scene of the virtual space set scene information of multiple items to be displayed according to an embodiment of the present disclosure is schematically shown.
[0162] As Figure 7 shown, Embodiment 700 includes the item information 701 of multiple items to be displayed, such as item information 701-1…. The item information 701 of multiple items to be displayed and the virtual image information 702 of the virtual image information are input into the large model M1 for generating the set scene information. The large model M1 for generating the set scene information can process the item information 701 of multiple items to be displayed to generate multiple differential information 703. For example, the differential information 703-1 corresponding to the item information 701-1…. After determining the normalization information 704 according to the multiple differential information 703, according to the normalization information, the virtual image information 702, and the single item information, such as the item information 701-1, multiple initial components of the item information 701-1 are generated, such as initial components 705-11…. According to the item information 701-1, the initial components 705-11 are updated to obtain the components 706-11. For the item information 701-1, according to multiple components, such as components 706-11…, the corresponding virtual space set scene sub-information 707-1 can be generated. The generation processes of other item information and components are similar and will not be elaborated here.
[0163] According to an embodiment of the present disclosure, a large model for generating scene information is called. According to multiple item information and virtual image information, virtual space scene information is generated, including: determining target items to be displayed from multiple items to be displayed according to the display levels of the multiple items to be displayed; calling the large model for generating scene information, and generating virtual space scene information according to the item information and virtual image information of the target items to be displayed.
[0164] The display level is used to represent the importance of multiple items to be displayed in the virtual space. In one embodiment, the display level can be determined through interaction with the user. For example, in an interaction operation, the user can select an item to be displayed with the highest display level through a selection operation; or, set the display levels of multiple items to be displayed through a selection operation. In another embodiment, it can also be determined according to the time when the item information of the item to be displayed is input. For example, the earlier the time when the item information of the item to be displayed is input, the higher the display level; conversely, the lower.
[0165] In one embodiment, the item to be displayed with the highest display level can be used as the target item to be displayed. The operation of calling the large model for generating scene information to generate virtual space scene information is as above and will not be elaborated here.
[0166] In the embodiments of the present disclosure, for a scene with multiple items to be displayed, the target item to be displayed is determined from the multiple items to be displayed through the display levels of the items to be displayed, and the large model for generating scene information is called. According to the item information and virtual image information of the target item to be displayed, virtual space scene information is generated, that is, taking the virtual space scene information of the target item to be displayed as the standard, it can avoid the scenes of multiple items to be displayed in the same virtual space from being too fragmented and improve the user experience.
[0167] In the embodiments of the present disclosure, the information of each component in the virtual space scene information can be stored as an independent object. The component contains descriptive metadata, such as component type, use, size, theme style, applicable scenario, etc. This can help the large model understand how to adapt the component to the virtual space scene information of a specific item. The data format of the descriptive metadata includes the definition and range of adjustable parameters, such as numerical values (such as size, position), enumerations (such as color scheme, style theme), or boolean values (such as whether to enable a certain special effect).
[0168] In addition, since the information of each component can be stored as an independent object, after generating the virtual space for displaying the item to be displayed, it also supports the user to adjust the attribute information of the component through interaction with the component in the virtual space. In one embodiment, the position and size of the component can be adjusted; the adjusted position and size by the user can also be used as feedback information to further optimize the large model for generating scene information.
[0169] In a specific embodiment, a large model for generating scene information can be trained using training samples (including sample item information and sample virtual image information), which may include: inputting the sample item information and sample virtual image information into the large model for generating scene information to be trained to obtain an output result. Comparing the output result with the label of the training sample, and adjusting the model parameters of the large model based on the comparison result, so that the output result gradually approaches the training sample, and finally obtaining the trained large model for generating scene information.
[0170] The training samples can be diverse sample item information, including item pictures, descriptions, comments, etc. These sample item information are used to help the large model for generating scene information (hereinafter referred to as the large model) understand the attributes and characteristics of the items. In addition, virtual space samples of various styles and themes need to be collected to ensure coverage of different decoration styles, color combinations, and layout designs. For example, the virtual space samples can be live broadcast scene samples, and these virtual space samples can be used as labels.
[0171] For example, the output result and the label can be input into a loss function to obtain a loss value. Adjusting the model parameters of the large model based on the loss value until the loss value converges. The type of the loss function is not limited. For example, it can be a cross-entropy loss function. The model parameters can include the learning rate, batch size, etc. The design of the loss function is to make the generated virtual space scene information match the item and the scene style consistent.
[0172] In addition, during the training phase, data augmentation can also be performed on the training samples to increase the diversity of the training samples and help the large model better generalize to unseen items and scenes. Or, user feedback and the virtual space scene information generated in actual applications can be used to continuously update and optimize the large model. Or, quantitative metrics and / or qualitative metrics can be used to evaluate the effect of the large model, and the large model can be continuously updated and optimized based on the evaluation results. For example, the quantitative metrics can be the quality score of the generated background image, style matching degree, etc., and the qualitative evaluation can be determined based on user evaluations.
[0173] Exemplarily, the sample virtual image information can include sample pose information. The sample pose information and sample item information can be input into the large model for generating scene information to be trained, and the loss value of the loss function can include at least one of the following: the loss value between the sample initial components generated according to the sample item information and the sample pose information and the label, the loss value between the sample components obtained by updating the sample initial components according to the sample item information and the label, and the loss value between the sample virtual space scene information generated according to the sample components and the label. Figure 8 The block diagram of the virtual space generation device according to an embodiment of the present disclosure is schematically shown.
[0174] As Figure 8As shown, the virtual space generation device includes an acquisition module 810, a first generation module 820, and a third generation module 830.
[0175] The acquisition module 810 is configured to acquire the item information of the item to be displayed and the virtual image information for displaying the item to be displayed in response to an interaction operation.
[0176] The first generation module 820 is configured to call a large model generated from the set design information, and generate virtual space set design information according to the item information and the virtual image information, wherein the virtual space set design information includes at least one component, and the number and display content of the components match the item information.
[0177] The second generation module 830 is configured to generate a virtual space for displaying the item to be displayed according to the virtual space set design information and the virtual image corresponding to the virtual image information.
[0178] According to an embodiment of the present disclosure, the virtual image information includes the pose information of the virtual image; the first generation module 820 includes: calling a large model generated from the set design information,
[0179] The first generation sub-module is configured to generate at least one initial component according to the item information and the pose information, wherein the number of the initial components matches the item information.
[0180] The update sub-module is configured to update the initial component according to the item information to obtain the component.
[0181] The second generation sub-module is configured to generate virtual space set design information according to at least one component.
[0182] According to an embodiment of the present disclosure, the first generation sub-module 820 includes: a pose estimation unit configured to perform pose estimation on the virtual image according to the pose information to obtain virtual image position information;
[0183] The first generation unit is configured to generate at least one initial component that matches the sub-item information of the item in each display dimension according to the item information, the virtual image position information, and the item type.
[0184] According to an embodiment of the present disclosure, the first generation unit includes: a first generation sub-unit configured to generate at least one initial component that matches the item information in an area different from the virtual image position information according to the item information and the pose information.
[0185] The adjustment unit is configured to adjust the positions between at least one initial component until at least one initial component that meets the predetermined position condition is generated when the positions between at least one initial component do not meet the predetermined position condition.
[0186] According to an embodiment of the present disclosure, the first generation sub-module 820 includes:
[0187] An extraction unit for extracting item sub-information of at least one display dimension and the item type of the item to be displayed from the item information.
[0188] A second generation unit for generating initial components matching the item sub-information of each display dimension according to the pose information and the item type.
[0189] According to an embodiment of the present disclosure, the second generation unit includes a first generation sub-unit for generating initial components matching the item sub-information of each display dimension according to the pose information and a candidate item type similar to the item type in the component library when the component library associated with the scene information generation large model does not include the item type.
[0190] A second generation sub-unit for determining candidate item sub-information of at least one candidate display dimension according to the evaluation information included in the item information when the component library does not include the item sub-information of the display dimension; and generating initial components matching the candidate item sub-information of each candidate display dimension according to the pose information and the item type.
[0191] According to an embodiment of the present disclosure, the first generation sub-module 820 further includes a background determination unit for determining at least one historical background image according to the pose information and the item type, where the historical background image includes background images generated using the same pose information and the same item type.
[0192] A background generation unit for generating a background image matching the item information according to at least one historical background image and the item information.
[0193] According to an embodiment of the present disclosure, the item information includes at least one item image information; the first generation module 820 further includes:
[0194] A first color system determination sub-module for determining target item image information from at least one item image information.
[0195] A second color system determination sub-module for determining the background color system from multiple colors according to the proportion of multiple colors in the target item image information.
[0196] A third generation sub-module for generating initial components matching the item sub-information of each display dimension according to the pose information, the item type, and the background color system.
[0197] According to an embodiment of the present disclosure, the second color system determination sub-module includes: a color system determination unit for determining the background color system according to the proportion of multiple colors in the background image and the proportion of multiple colors in the target item image information.
[0198] According to an embodiment of the present disclosure, the virtual space generation device 800 further includes:
[0199] A first custom information acquisition module, configured to acquire custom information, where the custom information includes setting style information and / or a custom color system.
[0200] A third generation module, configured to call the large model for generating setting information, and generate virtual space setting information according to the custom information, item information, and virtual image information.
[0201] According to an embodiment of the present disclosure, the virtual space generation device 800 further includes:
[0202] A second custom information acquisition module, configured to acquire custom information in response to an interaction operation on the virtual space setting information.
[0203] A fourth generation module, configured to call the large model for generating setting information, and update the virtual space setting information according to the custom information to obtain new virtual space setting information.
[0204] According to an embodiment of the present disclosure, the virtual space generation device 800 further includes:
[0205] A fifth generation module, configured to, when acquiring item information and virtual image information of multiple items to be displayed, call the large model for generating setting information, and generate virtual space setting information according to the multiple item information and virtual image information; wherein the virtual space setting information includes virtual space setting sub-information matching each item to be displayed, each virtual space setting sub-information includes at least one component matching the item information of each item to be displayed, and the multiple virtual space setting sub-information includes at least one background information.
[0206] According to an embodiment of the present disclosure, the fifth generation module includes:
[0207] A differentiation determination sub-module, configured to call the large model for generating setting information, and determine the item type and / or setting color system of each item to be displayed from the item information of each item to be displayed.
[0208] A normalization determination sub-module, configured to determine a target item type and / or a target setting color system according to the number of items to be displayed belonging to the same item type and / or belonging to the same setting color system, and determine the target item type and / or the target setting color system as normalization information.
[0209] A fourth generation sub-module, configured to generate virtual space setting information according to the normalization information, the multiple item information, and the virtual image information.
[0210] According to an embodiment of the present disclosure, the fifth generation module further includes:
[0211] An item determination sub-module, configured to determine a target item to be displayed from multiple items to be displayed according to the display levels of the multiple items to be displayed.
[0212] A fifth generation sub-module, configured to call a large model for generating scene information, and generate virtual space scene information according to the item information and virtual image information of the target item to be displayed.
[0213] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0214] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as above.
[0215] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as above.
[0216] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as above when executed by a processor.
[0217] Figure 9 A structural block diagram of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure is schematically shown.
[0218] In an embodiment of the present disclosure, inspired by the von Neumann architecture in modern computer theory, as Figure 9 shown, the AI intelligent agent 900 may include five core modules: an input module 910, a processing module 920, and an output module 930.
[0219] The input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (such as a user or an external environment), and converting it into a format that the AI intelligent agent 900 can understand and process. The input module 910 is the primary link for the AI intelligent agent 900 to interact with the outside world, enabling the AI intelligent agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0220] In the example, the input module 910 may input the item information, virtual image information, item image information, pose information, etc. described above.
[0221] In the example, the processing module 920 is the core support for the AI agent 900 to handle complex tasks. The processing module 920 is used to determine the target task according to the input information received by the input module 910, determine the scene information to generate a large model according to the target task, and the virtual space generation method described above can be executed by calling the large model generated by the scene information.
[0222] In the example, the output module 930 can output the virtual space scene information described above.
[0223] In the example, the processing module 920 may include a control unit 921, a storage unit 922, and an arithmetic unit 923.
[0224] During operation, the control unit 921 will continuously interact with the storage unit 922, the arithmetic unit 923, and / or the output module 930. However, it should be noted that in the embodiments of the present disclosure, the control unit 921 acts as a single initiator to initiate communication with the storage unit 922, the arithmetic unit 923, and / or the output module 930, and there is no communication coupling between the storage unit 922, the arithmetic unit 923, and the output module 930.
[0225] In the example, the performance of the control unit 921 can be closely related to the large model on which the AI agent 900 is based. To give full play to the capabilities of the large model, the internal structure of the control unit 921 can be designed to be highly configurable and extensible to cope with various different types of tasks and requirements in real scenarios, etc.
[0226] The storage unit 922 can be responsible for memorizing information such as historical conversations and event streams. Information such as item information, virtual image information, prompt information, differential information, and normalization information can be included in the storage unit 922.
[0227] The arithmetic unit 923 can be regarded as a predefined tool library. Controls for rendering, calculation modules, etc. as described above can be included in the arithmetic unit 923.
[0228] In the example, after the AI agent 900 obtains the item information and the virtual image information, the AI agent 900 can use the item information and the virtual image information to determine the target task, and determine the scene information to generate a large model according to the target task. The item information and the virtual image information can be stored in the storage unit 922 of the processing module 920. The control unit 921 can call the large model to obtain the input information from the storage unit 922 and process it to generate the virtual space scene information. The control unit 921 calls the rendering control in the arithmetic module 923 to render the virtual space scene information, or renders the virtual space according to the virtual space scene information and the virtual image. Then, the control unit 921 controls the virtual space scene information or the virtual space to be transmitted to the output module 930.
[0229] The AI intelligent agent 900 according to an embodiment of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.
[0230] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0231] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0232] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0233] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0234] Figure 10 A block diagram of an electronic device suitable for implementing a virtual space generation method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0235] As Figure 10 shown, the electronic device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0236] Multiple components in the electronic device 1000 are connected to the input / output (I / O) interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0237] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the virtual space generation method. For example, in some embodiments, the virtual space generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the virtual space generation method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured as the virtual space generation method in any other suitable way (e.g., by means of firmware).
[0238] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0239] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0240] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0241] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0242] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0243] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0244] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0245] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a virtual space, comprising: In response to an interaction operation, obtaining item information of an item to be displayed and virtual image information for displaying the virtual image of the item to be displayed; Invoking a large model for generating scene information, and generating virtual space scene information according to the item information and the virtual image information, wherein the virtual space scene information includes at least one component, and the number and display content of the components match the item information; And Generating a virtual space for displaying the item to be displayed according to the virtual space scene information and a virtual image corresponding to the virtual image information.
2. The method according to claim 1, wherein The virtual image information includes pose information of the virtual image; the step of invoking the large model for generating scene information and generating virtual space scene information according to the item information and the virtual image information includes: invoking the large model for generating scene information, Generating at least one initial component according to the item information and the pose information, wherein the number of the initial components matches the item information; Updating the initial components according to the item information to obtain the components; and Generating the virtual space scene information according to at least one of the components.
3. The method according to claim 2, wherein The step of generating at least one initial component that matches the item information according to the item information and the pose information includes: Performing pose estimation on the virtual image according to the pose information to obtain virtual image position information; and generating at least one of the initial components that matches the item information according to the item information, the virtual image position information, and the pose information.
4. The method according to claim 3, wherein, The step of generating at least one of the initial components according to the item information, the virtual image position information, and the pose information includes: Generating at least one initial component that matches the item information according to the item information and the pose information in a region different from the virtual image position information; In the case where the positions between at least one of the initial components do not meet a predetermined position condition, adjusting the positions between at least one of the initial components until at least one initial component that meets the predetermined position condition is generated.
5. The method according to any one of claims 2 to 4, wherein, The step of generating at least one initial component that matches the item information according to the item information and the pose information includes: Extracting item sub-information of at least one display dimension and the item type of the item to be displayed from the item information; and Generating the initial components that match the item sub-information of each display dimension according to the pose information and the item type.
6. The method according to claim 5, wherein The step of generating the initial components that match the item sub-information of each display dimension according to the pose information and the item type includes: In the case where the component library associated with the large model for generating scene information does not include the item type, generating the initial components that match the item sub-information of each display dimension according to the pose information and a candidate item type similar to the item type in the component library; In the case that the item sub - information of the display dimension is not included in the component library, determine the candidate item sub - information of at least one candidate display dimension according to the evaluation information included in the item information; and generate the initial components that match the candidate item sub - information of each candidate display dimension according to the pose information and the item type.
7. The method according to claim 5, wherein The initial component further includes a background image of the virtual space, and the method further includes: Determine at least one historical background image according to the pose information and the item type, where the historical background image includes background images generated using the same pose information and the same item type; Generate a background image that matches the item information according to at least one of the historical background images and the item information.
8. The method according to any one of claims 2 to 7, wherein, The item information includes at least one item image information; the step of generating the initial components that match the item sub - information of each display dimension according to the pose information and the item type includes: Determine the target item image information from at least one of the item image information; Determine the background color system from multiple colors according to the proportion of multiple colors in the target item image information; and Generate the initial components that match the item sub - information of each display dimension according to the pose information, the item type, and the background color system.
9. The method according to claim 8, wherein, The step of determining the background color system according to the proportion of multiple colors in the target item image information includes: Determine the background color system according to the proportion of multiple colors in the background image and the proportion of multiple colors in the target item image information.
10. The method according to any one of claims 1 - 9 further includes: Obtain custom information, where the custom information includes scenery style information and / or custom color system; and Call the scenery information generation large - model, and generate the virtual space scenery information according to the custom information, the item information, and the virtual image information.
11. The method according to any one of claims 1 - 10 further includes: In response to an interaction operation on the virtual space scenery information, obtain the custom information; Call the scenery information generation large - model, and update the virtual space scenery information according to the custom information to obtain the new virtual space scenery information.
12. The method according to any one of claims 1 - 11 further includes: In the case of obtaining the item information of multiple items to be displayed and the virtual image information, call the scenery information generation large - model, and generate the virtual space scenery information according to the multiple item information and the virtual image information; where the virtual space scenery information includes virtual space scenery sub - information that matches each item to be displayed, and each virtual space scenery sub - information includes at least one component that matches the item information of each item to be displayed, and the multiple virtual space scenery sub - information includes at least one background information.
13. The method according to claim 12, wherein, The step of calling the scenery information generation large - model and generating the virtual space scenery information according to the multiple item information and the virtual image information includes: Call the large model generated from the set scene information, and determine the item type and / or the set color system of each item to be displayed from the item information of each item to be displayed; Determine the target item type and / or the target set color system according to the number of items to be displayed belonging to the same item type and / or belonging to the same set color system, and determine the target item type and / or the target set color system as the normalized information; and Generate the virtual space set scene information according to the normalized information, the multiple item information, and the virtual image information.
14. The method according to claim 13, wherein, The step of calling the large model generated from the set scene information to generate the virtual space set scene information according to the multiple item information and the virtual image information includes: Determine the target item to be displayed from the multiple items to be displayed according to the display levels of the multiple items to be displayed; Call the large model generated from the set scene information, and generate the virtual space set scene information according to the item information of the target item to be displayed and the virtual image information.
15. A virtual space generation device, comprising: An acquisition module, configured to acquire the item information of the item to be displayed and the virtual image information for displaying the item to be displayed in response to an interaction operation; A first generation module, configured to call the large model generated from the set scene information, and generate the virtual space set scene information according to the item information and the virtual image information, wherein the virtual space set scene information includes at least one component, and the number and display content of the components match the item information; And A second generation module, configured to generate a virtual space for displaying the item to be displayed according to the virtual space set scene information and the virtual image corresponding to the virtual image information.
16. An intelligent agent, comprising: An input module, configured to receive input information; A processing module, configured to determine a target task according to the input information received by the input module, determine a large model according to the target task, and execute the method according to any one of claims 1 to 14 by calling the large model to obtain output information; An output module, configured to output the output information obtained by the processing module.
17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 14.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 14.
19. A computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 14.