Method and device for controlling virtual digital human to realize house watching
By obtaining the semantic map and virtual camera pose trajectory of the three-dimensional model of the house, combining map coding and motion control network, the virtual camera and digital people are controlled to interactively display, the problem of poor interactiveness of online house viewing is solved and user satisfaction is improved.
Patent Information
- Application Number
- CN202510193687.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing technology, online house viewing has poor interactivity, which affects users' house viewing experience and satisfaction.
By obtaining the semantic map of the three-dimensional model of the target house, the pose trajectory of the virtual camera and the action sequence of the virtual digital person, combining the map encoding network, the pose planning network and the action control network, the virtual camera and the digital person are controlled to interact and display, and provide house viewing services.
It enhances the interactiveness of online houses and improves user satisfaction and house viewing experience.
Smart Images

Figure CN120339468A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method and device for controlling a virtual digital human to achieve house viewing. Background Art
[0002] House viewing is a key link in real estate leasing and sales, aiming to enable potential customers to have a full and comprehensive understanding and experience of the internal situation of the candidate houses, which has a decisive impact on the decision-making of renting or purchasing a house. The existing house viewing mainly involves, after the user initially selects through online text introductions, real-shot pictures, floor plans, and AR three-dimensional models, the house intermediary personnel go to the house on-site for on-site guidance, explanation, and Q&A offline. Although the online multimedia house introductions are easily accessible, they lack interactivity. Although the offline intermediary personnel's house viewing is intuitive and comprehensive, it consumes time and physical strength, affecting the user's house viewing experience and ultimate satisfaction. Therefore, there is an urgent need for a house viewing interaction method with both online usability and real-time interactivity to solve the above problems. Summary of the Invention
[0003] In view of this, embodiments of this application provide a method, device, electronic device, and computer-readable storage medium for controlling a virtual digital human to achieve house viewing, so as to solve the problem of poor interactivity in online house viewing in the prior art.
[0004] In a first aspect of the embodiments of this application, a method for controlling a virtual digital human to achieve house viewing is provided, including: obtaining a house semantic map of a three-dimensional model of a target house, a pose trajectory of a virtual camera within a preset time range, and an action sequence of the virtual digital human within a preset time range, where the virtual camera is set inside the three-dimensional model of the house; extracting global features of the house semantic map, and determining an absolute pose adjustment action and an incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory; determining a digital human adjustment action of the virtual digital human based on the action sequence, the global features, and the pose trajectory; controlling the virtual digital human according to the digital human adjustment action, and controlling the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action, so as to provide a house viewing service for a user.
[0005] In the second aspect of the embodiments of the present application, a device for controlling a virtual digital human to realize a house viewing is provided, including: an acquisition module configured to acquire a house semantic map of a three-dimensional model of a target house, a pose trajectory of a virtual camera within a preset time range, and an action sequence of the virtual digital human within a preset time range, wherein the virtual camera is arranged inside the three-dimensional model of the house; a pose planning module configured to extract global features of the house semantic map and determine an absolute pose adjustment action and an incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory; an action adjustment module configured to determine a digital human adjustment action of the virtual digital human based on the action sequence, the global features, and the pose trajectory; and a control module configured to control the virtual digital human according to the digital human adjustment action and control the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action to provide a house viewing service for a user.
[0006] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0007] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0008] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: acquiring a house semantic map of a three-dimensional model of a target house, a pose trajectory of a virtual camera within a preset time range, and an action sequence of the virtual digital human within a preset time range, wherein the virtual camera is arranged inside the three-dimensional model of the house; extracting global features of the house semantic map and determining an absolute pose adjustment action and an incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory; determining a digital human adjustment action of the virtual digital human based on the action sequence, the global features, and the pose trajectory; controlling the virtual digital human according to the digital human adjustment action and controlling the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action to provide a house viewing service for a user. By adopting the above technical means, the problem of poor interactivity in online house viewing in the prior art can be solved, thereby enhancing the interactivity of online house viewing and improving user satisfaction. Description of the Drawings
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0010] Figure 1 It is a schematic flowchart of a method for controlling a virtual digital human to realize house viewing provided by an embodiment of the present application;
[0011] Figure 2 It is a schematic flowchart of another method for controlling a virtual digital human to realize house viewing provided by an embodiment of the present application;
[0012] Figure 3 It is a schematic structural diagram of a device for controlling a virtual digital human to realize house viewing provided by an embodiment of the present application;
[0013] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0014] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0015] A method and a device for controlling a virtual digital human to realize house viewing according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0016] Figure 1 It is a schematic flowchart of a method for controlling a virtual digital human to realize house viewing provided by an embodiment of the present application. Figure 1 The method for controlling a virtual digital human to realize house viewing can be executed by a computer or a server, or software on a computer or a server. As Figure 1 shown, the method for controlling a virtual digital human to realize house viewing includes:
[0017] S101, obtaining a house semantic map of a three-dimensional model of a target house, a pose trajectory of a virtual camera within a preset time range, and an action sequence of the virtual digital human within a preset time range, wherein the virtual camera is arranged inside the three-dimensional model of the house;
[0018] S102, extracting global features of the house semantic map, and determining an absolute pose adjustment action and an incremental pose adjustment action of the virtual camera according to the global features and the pose trajectory;
[0019] S103, determining a digital human adjustment action of the virtual digital human according to the action sequence, the global features, and the pose trajectory;
[0020] S104, control the virtual digital human according to the adjusted actions of the digital human, and control the virtual camera according to the actions of absolute pose adjustment and incremental pose adjustment, so as to provide users with house viewing services.
[0021] The house semantic map is a map of the three-dimensional model of the house that can be understood by the neural network model. The virtual camera is set inside the three-dimensional model of the house. By adjusting the pose of the virtual camera, different positions of the three-dimensional model of the house can be displayed. The virtual digital human is an intelligent robot for house viewing. The virtual digital human interacts with the user according to the adjusted actions of the digital human, and the virtual camera adjusts its own pose according to the actions of absolute pose adjustment and incremental pose adjustment, so as to capture different positions of the three-dimensional model of the house.
[0022] According to the technical solution provided by the embodiment of the present application, obtain the house semantic map of the three-dimensional model of the target house, the pose trajectory of the virtual camera within a preset time range, and the action sequence of the virtual digital human within a preset time range, wherein the virtual camera is set inside the three-dimensional model of the house; extract the global features of the house semantic map, and determine the actions of absolute pose adjustment and incremental pose adjustment of the virtual camera according to the global features and the pose trajectory; determine the adjusted actions of the digital human of the virtual digital human according to the action sequence, the global features and the pose trajectory; control the virtual digital human according to the adjusted actions of the digital human, and control the virtual camera according to the actions of absolute pose adjustment and incremental pose adjustment, so as to provide users with house viewing services. By adopting the above technical means, the problem of poor interactivity in online house viewing in the prior art can be solved, thereby enhancing the interactivity of online house viewing and improving user satisfaction.
[0023] Further, extracting the global features of the house semantic map includes: taking the semantic annotations and position encodings of K landmark points in the house semantic map as tokens and inputting them into the map encoding network, and outputting K tokens, where the K tokens represent the local semantics and position information of the K landmark points, and the map encoding network is an architecture of Transformer; obtaining the global features by cascading the K tokens.
[0024] The semantic annotation of the landmark point is the text annotation content of the landmark point, and the position encoding of the landmark point is the position of the landmark point. The map encoding network encodes the semantic annotation and position encoding of the landmark point to obtain the local semantics and position information. The local semantics and position information form tokens, and all the tokens are cascaded as the global features. K is a positive integer.
[0025] Determine the absolute pose adjustment action and the incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory, including: the pose trajectory contains multiple poses; determine the tokens corresponding to the landmark points captured by the virtual camera at each pose from the global features; after embedding the global features as guiding information into the pose planning network, input each pose and the token corresponding to this pose into the pose planning network, and output the absolute pose adjustment action and the incremental pose adjustment action, where the pose planning network is an architecture of a long short-term memory network.
[0026] When the pose planning network learns the global features, determine the absolute pose adjustment action and the incremental pose adjustment action based on each pose and the token corresponding to this pose. The absolute pose adjustment action is to select one from the preset N virtual camera viewing poses, and the incremental pose adjustment action is to perform relative movements of up, down, left, right, front, and back and change the orientation of up, down, left, and right based on the current virtual camera viewing pose.
[0027] Determine the digital human adjustment action of the virtual digital human based on the action sequence, the global features, and the pose trajectory, including: the pose trajectory contains multiple poses; determine the tokens corresponding to the landmark points captured by the virtual camera at each pose from the global features; after embedding the global features as guiding information into the action control network, input the action sequence, each pose, and the token corresponding to this pose into the action control network, and output the digital human adjustment action, where the action control network is an architecture of a long short-term memory network.
[0028] When the action control network learns the global features, determine the digital human adjustment action based on the action sequence, each pose, and the token corresponding to this pose. The digital human adjustment actions include, but are not limited to, actions such as welcome, nod, look at, point at, guide, wave, etc., and expressions such as smile, speak, focus, joy, etc.
[0029] After controlling the virtual digital human according to the digital human adjustment action and controlling the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action to provide the user with a house viewing service, the method further includes: calculating the time efficiency score, the process evaluation score, the final satisfaction score, and the interaction enthusiasm score of the house viewing service; performing a weighted sum of the time efficiency score, the process evaluation score, the final satisfaction score, and the interaction enthusiasm score to obtain a reward value; optimizing the model parameters of the map encoding network or the pose planning network or the action control network according to the reward value to perform reinforcement learning training on the map encoding network or the pose planning network or the action control network.
[0030] Calculate the time efficiency score based on the absolute pose adjustment action and the incremental pose adjustment action. Among them, the more the number of absolute pose adjustment actions and incremental pose adjustment actions, the lower the time efficiency score; calculate the process evaluation score based on the user's evaluation of each landmark point of the target house; calculate the final satisfaction score based on the user's selection intention evaluation and the benchmark intention evaluation; calculate the interaction enthusiasm score based on the number of user questions, the average character length of the questions, and the question frequency distribution.
[0031] The evaluation score can be calculated according to the user's expression when viewing each landmark point of the target house, or the user's direct evaluation of each landmark point. The average of the evaluation scores of all landmark points is used as the process evaluation score.
[0032] Obtain the preset prompt words, the house information of the target house, and the user's input information; based on the preset prompt words, house information, and input information, use the large language model to generate the answer information for the input information.
[0033] For example, the preset prompt words are "You are now a real estate agent and start showing a house to a client. The prepared house situation information is as follows. The house information includes the house type information, decoration information, home appliance information, surrounding environment, rental or purchase process, etc. of the target house. The input information is the user's question. The large language model generates the answer information for the input information based on the preset prompt words and the house information.
[0034] Figure 2 It is a schematic flowchart of another method for controlling a virtual digital human to realize house viewing provided by an embodiment of the present application. As Figure 2 shown, the method includes:
[0035] S201, use the semantic annotation and position encoding of K landmark points in the house semantic map as tokens to input into the map encoding network, and output K tokens;
[0036] S202, cascade the features of the K tokens to obtain the global feature.
[0037] S203, determine the tokens corresponding to the landmark points captured by the virtual camera in each pose from the global feature;
[0038] S204, after embedding the global feature as the guiding information into the pose planning network, input each pose and the token corresponding to this pose into the pose planning network, and output the absolute pose adjustment action and the incremental pose adjustment action;
[0039] S205, after embedding the global feature as the guiding information into the action control network, input the action sequence, each pose, and the token corresponding to this pose into the action control network, and output the digital human adjustment action;
[0040] S206. Based on the preset prompt words, housing information, and input information, use a large language model to generate response information for the input information;
[0041] S207. Display the response information, control the virtual digital human according to the digital human's adjusted actions, and control the virtual camera according to the absolute adjustment of the pose and the incremental adjustment of the pose to provide the user with a housing viewing service.
[0042] An embodiment of the present application provides a virtual digital human interaction system. The virtual digital human interaction system is composed of a virtual digital human generation module, a multimodal dialogue module, a housing three-dimensional display module, a language large model interface module, an intelligent planning and control module, and a visual interaction interface.
[0043] The virtual digital human generation module is used to generate virtual digital humans with actions and expressions. The action library of the virtual digital human includes but is not limited to welcome, nod, look at, point to, guide, wave, etc. The expression library of the virtual digital human includes but is not limited to smile, speak, concentrate, joy, etc.
[0044] The multimodal dialogue module includes voice input, speech synthesis, text input, and text display, and is used to interact with the user in natural speech or language. If the user uses voice input, the input voice is converted into input text information. If the user uses text input, it is directly used as input text information. The output text information can be selected from the output text library or received as the generation result of the cloud language large model. When the output text information is fed back to the user, speech synthesis and text display are performed simultaneously, enabling the user to not only listen to a more convenient voice feedback but also view a more accurate text feedback. Each text segment in the output text library corresponds to a landmark point on the housing semantic map. When the virtual camera's field of view moves to a new landmark point, the corresponding text segment in the output text library is triggered and output.
[0045] The housing three-dimensional display module directly calls the reconstructed three-dimensional model of the housing interior, allowing the viewpoint and perspective of the virtual camera to be set, so as to display different orientations, positions, and ranges inside the housing. The viewpoint and perspective settings of the virtual camera have two modes: absolute adjustment and incremental adjustment. Absolute adjustment is to select one from the preset N virtual camera observation poses, and incremental adjustment is to perform relative movement up, down, left, right, forward, and backward and change the orientation up, down, left, and right based on the current virtual camera observation pose.
[0046] The language large model interface module is used to send text prompts to the cloud language large model and receive the generated text from the cloud language large model. When the language large model interface module is initialized for the first time, it first sends a formatted text to the cloud language large model, with the content "You are now a real estate agent. Now start showing the house to the client. The prepared house situation information is as follows:", followed by the description text of all the house situation information, including but not limited to house type information, decoration information, home appliance information, surrounding environment, rental or purchase process, etc. After initialization, when the text information obtained by the user through voice input or text input, the text information is sent to the cloud language large model as a text prompt and the generated text from the cloud language large model is received.
[0047] The intelligent planning and control module is used to plan the house viewing path and control the actions of the digital virtual human, with the goal of maximizing the user's house viewing efficiency and satisfaction.
[0048] Build a viewing strategy model based on a deep neural network. The input of the viewing strategy model is the house semantic map M and the virtual camera pose trajectory {Pt}, and the output is the virtual digital human action AV, the absolute adjustment action AA of the virtual camera pose, and the incremental adjustment action AI of the virtual camera pose.
[0049] The viewing strategy model has a house semantic map encoding network, which is implemented based on the Transformer architecture. The semantic annotations and two-dimensional position encodings of K landmark points in the house semantic map M are used as tokens and input into the Transformer network. The output K tokens {Tk} respectively represent the local semantic and position information of the K landmark points. After feature concatenation of the K tokens, a global house semantic map description TG is formed.
[0050] The viewing strategy model has a virtual camera pose planning network, which is implemented based on the long short-term memory network architecture. First, the virtual camera pose trajectory {Pt} within the historical T time range and the tokens Tk corresponding to the position indices of each Pt are input into the long short-term memory network, and the virtual camera pose adjustment actions AA,next and AI,next at the next moment are output. The global house semantic map description is embedded into the long short-term memory network as guiding information.
[0051] The viewing measurement model has a virtual digital human action control network, which is implemented based on the long short-term memory network architecture. First, the virtual digital human action sequence {AV,t}, the virtual camera pose trajectory {Pt} within the historical T time range, and the tokens Tk corresponding to the position indices of each Pt are input into the long short-term memory network, and the virtual digital human action AV at the next moment is output. The global house semantic map description is embedded into the long short-term memory network as guiding information.
[0052] The viewing strategy model is trained using a deep Q-network reinforcement learning framework. First, the viewing expert conducts offline house viewing demonstrations. The viewing expert can operate the viewpoint and perspective adjustment of the virtual camera while controlling the action types of the virtual digital human at various punctuation points. After collecting a certain amount of expert demonstration data, it is used to supervise the training of the viewing strategy model to form an initial viewing strategy, so as to improve the efficiency of subsequent reinforcement learning.
[0053] A house viewing reward function R is established, which is formed by the weighted sum of the time efficiency term Reff, the process evaluation term Rrate, the final satisfaction term Rsati, and the interaction enthusiasm term Rinter, that is, R = w1Reff + w2Rrate + w3Rsati + w4Rinter.
[0054] The time efficiency term Rtime is calculated using the total length of the virtual camera pose path, guiding the strategy to tend to improve the house viewing efficiency and avoid redundant paths.
[0055] The process evaluation term Rrate is realized using the average value of the increment of the viewing experience evaluation value of the user at various punctuation points, guiding the strategy to improve the user's experience during the house viewing process.
[0056] The final satisfaction term Rsati is realized using the deviation value between the evaluation of the user's selection intention for the house after viewing and the benchmark selection intention evaluation, guiding the strategy to increase the user's satisfaction and positive selection intention after viewing the house.
[0057] The interaction enthusiasm term Rinter is calculated using the number of questions and answers, the average character length of the questions and answers, and the question and answer frequency distribution between the user and the multi-modal interaction module, and a maximum limit is set, guiding the strategy to improve the user's interaction enthusiasm during the house viewing process while avoiding excessive interaction.
[0058] In the online reinforcement learning stage, the viewing strategy model and the greedy algorithm are used to select the output as the action AV of the virtual digital human, the absolute adjustment action AA of the virtual camera pose, and the incremental adjustment action AI of the virtual camera pose. After the system executes the above actions, the reward value is calculated according to the user feedback and state changes to obtain a new state, and the state, action, reward value, and new state are stored in the experience pool. A small batch of experiences is randomly sampled from the experience pool, the target Q value is calculated, and the parameters of the viewing strategy model are updated using gradient descent to minimize the difference between the predicted Q value and the target Q value.
[0059] The visual interaction interface displays the interior view of the house, the virtual digital human image, the text display box, and the user selection menu window within the field of view of the current virtual camera. The virtual digital human image is displayed on the left or right side of the visual interaction interface to leave enough screen space for the house view. When the virtual digital human makes a pointing gesture, the key parts of the interior view of the house are emphasized by highlighting or marking to guide the user's visual attention. The text display box is located at the lower edge of the visual interaction interface. The user selection menu window is pop-up, presenting several evaluation or decision options, and the window automatically closes after the user makes a selection.
[0060] The embodiments of the present application provide a high-efficiency, low-cost, and highly accessible housing viewing solution. By combining virtual digital humans and artificial intelligence technologies, it helps to improve the service level and user experience in the real estate market, and mainly has the following advantages: Through virtual digital human interaction, it provides a more user-friendly and interactive housing viewing experience. Users can communicate with the virtual digital human at any time to obtain information. The adjustment of the 3D house model and the virtual camera perspective allows users to more intuitively understand the layout and details of the house; The intelligent planning control module can automatically plan the optimal viewing path according to the user's preferences and the characteristics of the house, improving the viewing efficiency and user experience; The expression and action library of the virtual digital human can be adaptively adjusted according to the user's feedback to provide more personalized services. The language large model interface module can handle complex conversations and provide detailed and accurate housing information, enhancing the user's interactive experience.
[0061] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present application, which will not be elaborated one by one here.
[0062] The following are the device embodiments of the present application, which can be used to execute the method embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0063] Figure 3 It is a schematic diagram of a device for controlling a virtual digital human to achieve housing viewing provided by the embodiments of the present application. As Figure 3 shown, the device for controlling a virtual digital human to achieve housing viewing includes:
[0064] An acquisition module 301, configured to acquire the house semantic map of the house 3D model of the target house, the pose trajectory of the virtual camera within a preset time range, and the action sequence of the virtual digital human within a preset time range, wherein the virtual camera is set inside the house 3D model;
[0065] A pose planning module 302, configured to extract the global features of the house semantic map, and determine the absolute pose adjustment action and the incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory;
[0066] The action adjustment module 303 is configured to determine the digital human adjustment actions of the virtual digital human according to the action sequence, the global feature, and the pose trajectory.
[0067] The control module 304 is configured to control the virtual digital human according to the digital human adjustment actions, and control the virtual camera according to the absolute pose adjustment actions and the incremental pose adjustment actions, so as to provide the user with a house viewing service.
[0068] According to the technical solution provided by the embodiment of the present application, obtain the house semantic map of the three-dimensional model of the target house, the pose trajectory of the virtual camera within a preset time range, and the action sequence of the virtual digital human within a preset time range, wherein the virtual camera is set inside the three-dimensional model of the house; extract the global feature of the house semantic map, and determine the absolute pose adjustment action and the incremental pose adjustment action of the virtual camera according to the global feature and the pose trajectory; determine the digital human adjustment action of the virtual digital human according to the action sequence, the global feature, and the pose trajectory; control the virtual digital human according to the digital human adjustment action, and control the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action, so as to provide the user with a house viewing service. By adopting the above technical means, the problem of poor interactivity in the existing online house viewing can be solved, thereby enhancing the interactivity of the online house viewing and improving the user satisfaction.
[0069] In some embodiments, the pose planning module 302 is further configured to use the semantic annotation and position encoding of K landmark points in the house semantic map as tokens to input into the map encoding network, and output K tokens, where the K tokens represent the local semantics and position information of the K landmark points, and the map encoding network is of the Transformer architecture; perform feature concatenation on the K tokens to obtain the global feature.
[0070] The semantic annotation of the landmark point is the text annotation content of the landmark point, and the position encoding of the landmark point is the position of the landmark point. The map encoding network encodes the semantic annotation and position encoding of the landmark point to obtain the local semantics and position information. The local semantics and position information form tokens, and all tokens are concatenated as the global feature. K is a positive integer.
[0071] In some embodiments, the pose planning module 302 is further configured to the pose trajectory includes multiple poses; determine the tokens corresponding to the landmark points captured by the virtual camera at each pose from the global feature; after embedding the global feature as the guiding information into the pose planning network, input each pose and the token corresponding to the pose into the pose planning network, and output the absolute pose adjustment action and the incremental pose adjustment action, where the pose planning network is of the long short-term memory network architecture.
[0072] When the pose planning network has learned the global features, it determines the absolute pose adjustment action and the incremental pose adjustment action based on each pose and the token corresponding to that pose. The absolute pose adjustment action is to select one from the preset N virtual camera viewing poses, and the incremental pose adjustment action is to perform relative movements up, down, left, right, forward, and backward and change the orientation up, down, left, and right based on the current virtual camera viewing pose.
[0073] In some embodiments, the action adjustment module 303 is further configured such that the pose trajectory includes multiple poses; to determine the tokens corresponding to the landmark points captured by the virtual camera at each pose from the global features; after embedding the global features as guiding information into the action control network, to input the action sequence, each pose, and the token corresponding to that pose into the action control network, and output the digital human adjustment action, where the action control network is an architecture of a long short-term memory network.
[0074] When the action control network has learned the global features, it determines the digital human adjustment action based on the action sequence, each pose, and the token corresponding to that pose. The digital human adjustment actions include, but are not limited to, actions such as welcoming, nodding, looking at, pointing, guiding, waving, etc., as well as expressions such as smiling, speaking, concentrating, and being joyful.
[0075] In some embodiments, the control module 304 is further configured to calculate the time efficiency score, the process evaluation score, the final satisfaction score, and the interaction enthusiasm score of the house viewing service; to perform a weighted sum of the time efficiency score, the process evaluation score, the final satisfaction score, and the interaction enthusiasm score to obtain a reward value; and to optimize the model parameters of the map encoding network, the pose planning network, or the action control network based on the reward value to perform reinforcement learning training on the map encoding network, the pose planning network, or the action control network.
[0076] In some embodiments, the control module 304 is further configured to calculate the time efficiency score based on the absolute pose adjustment action and the incremental pose adjustment action, where the more the number of absolute pose adjustment actions and incremental pose adjustment actions, the lower the time efficiency score; to calculate the process evaluation score based on the user's evaluations of each landmark point of the target house; to calculate the final satisfaction score based on the user's selection intention evaluation and the benchmark intention evaluation; and to calculate the interaction enthusiasm score based on the number of user questions and answers, the average character length of the questions and answers, and the question and answer frequency distribution.
[0077] The evaluation score can be calculated based on the expressions of the user viewing each landmark point of the target house, or the user directly gives the evaluations of each landmark point, and the average of the evaluation scores of all landmark points is used as the process evaluation score.
[0078] In some embodiments, the control module 304 is further configured to obtain a preset prompt word, the housing information of the target house, and the input information of the user; and generate response information for the input information by using a large language model based on the preset prompt word, the housing information, and the input information.
[0079] For example, the preset prompt word is "You are now a real estate agent and start showing a client around a house. The prepared housing situation information is as follows. The housing information includes the house type information, decoration information, household appliance information, surrounding environment, rental or purchase process, etc. of the target house. The input information is the user's question. The large language model generates response information for the input information based on the preset prompt word and the housing information.
[0080] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0081] Figure 4 is a schematic diagram of the electronic device 4 provided by the embodiment of the present application. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above various method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.
[0082] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4, and do not constitute a limitation to the electronic device 4, and may include more or fewer components than those shown in the figure, or different components.
[0083] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0084] The memory 402 may be an internal storage unit of the electronic device 4, for example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk equipped on the electronic device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0085] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0086] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0087] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for controlling a virtual digital human to achieve house viewing, characterized in that, Including: Obtaining a house semantic map of a target house's three-dimensional model, the pose trajectory of a virtual camera within a preset time range, and the action sequence of a virtual digital human within the preset time range, where the virtual camera is set inside the house three-dimensional model; Extracting the global features of the house semantic map, and determining the absolute pose adjustment action and incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory; Determining the digital human adjustment action of the virtual digital human based on the action sequence, the global features, and the pose trajectory; Controlling the virtual digital human according to the digital human adjustment action, and controlling the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action to provide a house viewing service for the user.
2. The method according to claim 1, characterized in that, Extracting the global features of the house semantic map includes: Taking the semantic annotations and position encodings of K landmark points in the house semantic map as tokens and inputting them into a map encoding network, and outputting K tokens, where the K tokens represent the local semantics and position information of the K landmark points, and the map encoding network is an architecture of Transformer; Obtaining the global features by concatenating the features of the K tokens.
3. The method according to claim 1, characterized in that, Determining the absolute pose adjustment action and incremental pose adjustment action of the virtual camera based on the global features and the pose trajectory includes: The pose trajectory contains multiple poses; Determining the tokens corresponding to the landmark points captured by the virtual camera at each pose from the global features; After embedding the global features as guiding information into a pose planning network, inputting each pose and the token corresponding to this pose into the pose planning network, and outputting the absolute pose adjustment action and the incremental pose adjustment action, where the pose planning network is an architecture of long short-term memory network.
4. The method according to claim 1, wherein Determining the digital human adjustment action of the virtual digital human based on the action sequence, the global features, and the pose trajectory includes: The pose trajectory contains multiple poses; Determining the tokens corresponding to the landmark points captured by the virtual camera at each pose from the global features; After embedding the global features as guiding information into an action control network, inputting the action sequence, each pose, and the token corresponding to this pose into the action control network, and outputting the digital human adjustment action, where the action control network is an architecture of long short-term memory network.
5. The method according to claim 2 or 3 or 4, characterized in that, After controlling the virtual digital human according to the digital human adjustment action, and controlling the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action to provide a house viewing service for the user, the method further includes: Calculating the time efficiency score, process evaluation score, final satisfaction score, and interaction enthusiasm score of the house viewing service; Performing a weighted sum of the time efficiency score, the process evaluation score, the final satisfaction score, and the interaction enthusiasm score to obtain a reward value; Optimize the model parameters of the map encoding network or the pose planning network or the motion control network according to the reward value to perform reinforcement learning training on the map encoding network or the pose planning network or the motion control network.
6. The method according to claim 5, wherein It includes: Calculate the time efficiency score based on the absolute pose adjustment action and the incremental pose adjustment action, where the more the number of the absolute pose adjustment action and the incremental pose adjustment action, the lower the time efficiency score; Calculate the process evaluation score based on the user's evaluation of each landmark point of the target house; Calculate the final satisfaction score based on the user's selection intention evaluation and the benchmark intention evaluation; Calculate the interaction enthusiasm score based on the number of questions and answers of the user, the average character length of the questions and answers, and the question and answer frequency distribution.
7. The method according to claim 1, wherein The method further includes: Obtain a preset prompt word, the house information of the target house, and the input information of the user; According to the preset prompt word, the house information, and the input information, use a large language model to generate an answer information for the input information.
8. A device for controlling a virtual digital human to realize house viewing, characterized in that, It includes: An acquisition module, configured to acquire the house semantic map of the three-dimensional model of the target house, the pose trajectory of the virtual camera within a preset time range, and the action sequence of the virtual digital human within the preset time range, where the virtual camera is set inside the three-dimensional model of the house; A pose planning module, configured to extract the global features of the house semantic map, and determine the absolute pose adjustment action and the incremental pose adjustment action of the virtual camera according to the global features and the pose trajectory; An action adjustment module, configured to determine the digital human adjustment action of the virtual digital human according to the action sequence, the global features, and the pose trajectory; A control module, configured to control the virtual digital human according to the digital human adjustment action, and control the virtual camera according to the absolute pose adjustment action and the incremental pose adjustment action to provide the user with a house viewing service.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.