Interaction method, apparatus, device, and storage medium
By acquiring the intent characteristics of the target object, adjusting the image characteristics of the intelligent digital human and outputting response information, the problem of understanding and applying virtual fitting models has been solved, and multi-dimensional interaction and efficient user experience improvement have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
- Filing Date
- 2024-09-23
- Publication Date
- 2026-04-28
AI Technical Summary
Existing virtual try-on models are difficult to understand and apply, requiring a significant amount of time and practice to achieve the desired results.
By acquiring the intent characteristics of the target object, adjusting the image characteristics of the intelligent digital human, and outputting corresponding response information, multi-dimensional interaction is achieved, including intelligent dialogue, clothing recommendations, and product feature display.
It enriches the user experience, enhances user immersion and interaction, saves time and effort, and improves the effectiveness of virtual try-on.
Smart Images

Figure CN119579830B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of data processing and image processing, and in particular to the fields of artificial intelligence and large models. Background Technology
[0002] Currently, many e-commerce platforms have launched virtual try-on models in an attempt to gain a technological advantage in the e-commerce field. However, despite these models' technological breakthroughs, they present certain challenges in understanding and application. Therefore, achieving the ideal effect of virtual try-on still requires a significant amount of time and practice to explore and refine. Summary of the Invention
[0003] This disclosure provides an interaction method, apparatus, device, and storage medium.
[0004] According to one aspect of this disclosure, an interaction method is provided, comprising:
[0005] In response to a first control operation, a first intent feature of the first control operation is acquired, wherein the first intent feature is at least used to instruct the intelligent digital human to adjust its image features;
[0006] Based on the target trigger command indicated by the first intent feature, adjust the image features of the intelligent digital human to display the adjusted intelligent digital human; and,
[0007] Output the first response information corresponding to the first intent feature.
[0008] According to another aspect of this disclosure, an interactive device is provided, comprising:
[0009] A processing unit is configured to, in response to a first control operation, acquire a first intent feature of the first control operation, wherein the first intent feature is at least used to instruct the intelligent digital human to adjust its image features; adjust the image features of the intelligent digital human based on a target trigger instruction indicated by the first intent feature to obtain an adjusted intelligent digital human; and...
[0010] The output unit is used to display the adjusted intelligent digital human and output the first response information corresponding to the first intent feature.
[0011] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0012] At least one processor; and
[0013] The memory is communicatively connected to the at least one processor; wherein,
[0014] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0015] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0016] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0017] In this way, the disclosed solution can adjust the image features of the intelligent digital human according to the intent features of the target object, and at the same time output response information corresponding to the intent features. Thus, it realizes multi-dimensional interaction with the target object, thereby enriching and improving the user experience.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0019] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0020] Figure 1 This is a schematic diagram of the implementation flow of an interaction method according to an embodiment of this application. Figure 1 ;
[0021] Figure 2 This is a schematic diagram of the implementation flow of an interaction method according to an embodiment of this application. Figure 2 ;
[0022] Figures 3(a) to 3(e) This is a schematic diagram of an interactive scenario of an intelligent digital human according to an embodiment of this application. Figure 1 ;
[0023] Figure 4 This is a schematic flowchart of an interaction method according to an embodiment of this application;
[0024] Figures 5(a) to 5(d) This is a schematic diagram of an interactive scenario of an intelligent digital human according to an embodiment of this application. Figure 2 ;
[0025] Figure 6 This is a schematic diagram illustrating a comparison scenario of product feature information of the same product object according to an embodiment of this application;
[0026] Figures 7(a) and 7(b) are schematic flowcharts of an interaction method according to an embodiment of the present application in a specific embodiment;
[0027] Figure 8 This is a schematic diagram of the structure of an interactive device according to an embodiment of this application;
[0028] Figure 9 This is a block diagram of an electronic device used to implement the interaction method of the embodiments of this disclosure. Detailed Implementation
[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0030] In this document, the term "and / or" merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document indicates any combination of at least two of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this document refer to and distinguish between multiple similar technical terms, not to restrict the order or to limit there to only two. For example, "first feature" and "second feature" refer to two categories / two features; the first feature can be one or more, and the second feature can also be one or more.
[0031] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can still be practiced even without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0032] This disclosed solution provides a specific approach that enables intelligent digital humans to change their appearance features in real time, such as changing clothes, and to display the three-dimensional effect of clothing after try-on in multiple dimensions. Furthermore, by utilizing intelligent digital humans (e.g., hyper-realistic digital humans), this solution can also engage in intelligent dialogue and make intelligent recommendations (e.g., clothing and styling recommendations) while changing appearance features. It can even display product feature information for the same item from different recommendation platforms, thus enriching and enhancing the user experience.
[0033] Specifically, Figure 1 This is an illustrative flow diagram of an interaction method according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0034] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes:
[0035] Step S101: In response to the first control operation, obtain the first intent feature of the first control operation.
[0036] Here, the first intention feature is at least used to instruct the intelligent digital human to adjust its appearance features. For example, in one example, the appearance features include, but are not limited to, at least one of the following: facial features (such as expressions, movements, and makeup), hairstyle features (such as hair color and hairstyle), clothing, etc.
[0037] Furthermore, in one example, the first control operation may be specifically a voice control operation or a text control operation.
[0038] It should be noted that the intelligent digital human described in this disclosure can be specifically a 3-dimensional (D) intelligent digital human, or more specifically a 3D intelligent digital human.
[0039] Step S102: Based on the target trigger command indicated by the first intent feature, adjust the image features of the intelligent digital human to display the adjusted intelligent digital human.
[0040] Step S103: Output the first response information corresponding to the first intent feature.
[0041] For example, in one instance, the intelligent digital human can be used to output first response information corresponding to the first intention feature. Alternatively, a digitally enhanced intelligent human with adjusted appearance can be used to output first response information corresponding to the first intention feature.
[0042] Furthermore, in one example, the first response information can be specifically output via voice by the intelligent digital human. For instance, after receiving the first response information, it can be converted into response voice data and then played back using the intelligent digital human. Furthermore, to further enhance the user experience and display effect, the mouth movements of the intelligent digital human can be rendered based on the response voice data and displayed synchronously with the voice broadcast.
[0043] It should be noted that the first response information can also be displayed as text content. For example, if at least some areas of the display page other than the intelligent digital human are set up with text display boxes, the text content of the first response information can be displayed in the text display boxes.
[0044] In this way, the disclosed solution can adjust the image features of the intelligent digital human according to the intent features of the target object, and at the same time output response information corresponding to the intent features. Thus, it realizes multi-dimensional interaction with the target object, thereby enriching and improving the user experience.
[0045] For example, in the scenario of intelligent digital human try-on, the disclosed solution can effectively realize the change of clothes through intelligent digital human according to the needs (i.e., intent characteristics) of the target object, and then intuitively show the try-on effect of clothing (including clothes and accessories) on intelligent digital human. This process does not require the target object to try on the clothes in person, which effectively saves the target object's time and energy. Moreover, through the intelligent digital human changing clothes and interacting with the target object, it can also effectively reduce the target object's sense of detachment during the browsing process, so that the target object can obtain a more immersive browsing experience, thereby further improving the user experience.
[0046] In a specific example, when displaying the adjusted intelligent digital human, the adjusted intelligent digital human can be triggered to rotate around a preset axis for multi-dimensional display; or, when displaying the adjusted intelligent digital human, in response to a display operation (e.g., in response to a touch operation of a display button (e.g., a virtual button) on the display interface), the adjusted intelligent digital human can be triggered to rotate around a preset axis for multi-dimensional display. This further demonstrates the overall effect of the adjusted image features, thereby further enhancing the browsing experience.
[0047] Here, in one example, the multi-dimensional display can specifically refer to a 360-degree all-around display. For instance, in one example, while displaying the adjusted intelligent digital human, the intelligent digital human can also automatically rotate 360 degrees around a preset axis for a 360-degree display. This facilitates viewing the adjusted image of the intelligent digital human from all angles, thereby effectively improving the user experience.
[0048] Alternatively, in another example, when displaying the adjusted intelligent digital human, a 360-degree display instruction for the target object is obtained. Based on this 360-degree display instruction, the intelligent digital human is automatically triggered to rotate 360 degrees around a preset axis.
[0049] Thus, this disclosed solution enables observation of the details of the intelligent digital human's appearance from multiple angles, further enhancing the browsing experience. Moreover, multi-dimensional, such as 360-degree display, can enhance the immersion and interactive experience of the target audience, thereby significantly improving the realism of the user experience.
[0050] In a specific example of the scheme disclosed herein, the first intent feature can be obtained in the following manner; specifically, the above-described response to the first control operation, obtaining the first intent feature of the first control operation (e.g., step S101), specifically includes:
[0051] In response to the first control operation, the operation features of the first control operation (e.g., text data corresponding to the first control operation) are input into the second model to obtain the first intent feature of the first control operation.
[0052] For example, if the first control operation is a voice control operation, the voice data corresponding to the voice control operation can be converted into text data first, and the text data of the voice control operation can be input into the second model to obtain the first intent feature of the voice control operation.
[0053] Alternatively, if the first control operation is a text control operation, the text data corresponding to the text control operation can be directly input into the second model to obtain the first intent feature of the text control operation.
[0054] In other words, in this example, the second model can be used to understand intent, and then the first intent feature of the first control operation can be quickly inferred.
[0055] Here, the second model can be specifically a large language model (also known as a large model), or it can be other natural language processing models with content reasoning capabilities. This disclosure does not impose any specific restrictions on this.
[0056] In this way, the disclosed solution can use the model to gain a deep understanding of the control operations input by the target object, so as to accurately locate the needs of the target object and facilitate subsequent operations based on the understood needs, thus laying the foundation for improving the user experience.
[0057] Furthermore, in a specific example, the above-described response to a first control operation, inputting the operational features of the first control operation into a second model to obtain a first intent feature of the first control operation, specifically includes:
[0058] In response to a first control operation containing voice data, the first control operation is converted into target text data;
[0059] The target text data is input into the second model to obtain the first intent feature of the first control operation.
[0060] For example, in one instance, the target object inputs a piece of voice data to control the intelligent digital human to adjust its appearance features. In this case, speech-to-text technology can be used to convert the voice data input by the target object (that is, the voice data corresponding to the voice control operation, also known as control voice data) into target text data. The obtained target text data is then input into a large language model for intent understanding to obtain the target object's intent features (that is, the first intent features).
[0061] Here, the speech-to-text technology in this example can be specifically an acoustic model-based recognition technology, or it can be a language model-based recognition technology; this disclosure does not limit it in this way.
[0062] In this way, the disclosed solution can effectively respond to control voice data and convert the control voice data into text, and then use a large model for deep understanding and complex reasoning to obtain intent features. This significantly improves the efficiency and accuracy of information processing, realizes a deep understanding of user needs, and lays the foundation for providing more accurate feedback in the future.
[0063] Figure 2 This is an illustrative flow diagram of an interaction method according to an embodiment of this application. Figure 2 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 The methods shown can also be applied to this example, but the related content will not be elaborated further in this example.
[0064] Furthermore, the method includes at least a portion of the following: For example... Figure 2 As shown, it includes:
[0065] Step S201: In response to the first control operation, obtain the first intent feature of the first control operation.
[0066] Here, the first intention feature is used at least to instruct the intelligent digital human to adjust its image features.
[0067] It should be noted that the relevant content regarding the first intent feature can be found in the above example, and will not be repeated here.
[0068] Step S202: Based on the target trigger command indicated by the first intent feature, adjust the image features of the intelligent digital human to display the adjusted intelligent digital human.
[0069] Step S203: Generate target prompt words based on the first intent feature of the first control operation and the preset prompt word template.
[0070] For example, in one example, the preset prompt word template contains preset examples (such as question-answer pairs). In this case, based on the first intent feature and the preset examples of the preset prompt word template, a target prompt word that matches the first intent feature can be obtained (such as a question-answer pair that meets the requirements). In this way, the target prompt word can effectively guide the first model to perform reasoning and thus obtain the first response information that meets the expectations.
[0071] Step S204: Input the target prompt word into the first model to obtain the first response information.
[0072] Here, the first model can be specifically a large language model (also known as a large model), or it can be other natural language processing models with content reasoning capabilities. This disclosure does not impose any specific restrictions on this.
[0073] Step S205: Output the first response information corresponding to the first intent feature.
[0074] For example, the intelligent digital human can be used to output first response information corresponding to the first intention feature. Alternatively, the intelligent digital human with an adjusted appearance can be used to output first response information corresponding to the first intention feature. Specific examples can be found in the above description and will not be repeated here.
[0075] It should be noted that some of the above steps can be executed simultaneously, or the execution order can be changed. For example, the execution order of steps S202 and steps S203 to S205 can be executed simultaneously, or steps S203 to S205 can be executed first, and then step S202 can be executed, etc. This disclosure does not limit this.
[0076] Step S206: Obtain the first response information in response to the first response information.
[0077] Step S207: Input the first response information into the first model at least to obtain the second response information in response to the first response information.
[0078] Step S208: Output the second response information to perform intelligent question answering.
[0079] For example, a second response message can be output using an intelligent digital human. Or, further, a second response message can be output using an intelligent digital human with an adjusted appearance.
[0080] Furthermore, in one example, the second response information can be specifically output via voice by the intelligent digital human. For instance, after obtaining the second response information, it can be converted into response voice data and then played back using the intelligent digital human. Furthermore, to further enhance the user experience and display effect, the mouth movements of the intelligent digital human can be rendered based on the response voice data and displayed synchronously with the voice broadcast.
[0081] It should be noted that the second response information can also be displayed as text content. For example, if at least some areas of the display page other than the intelligent digital human are set up with text display boxes, the text content of the second response information can be displayed in the text display boxes.
[0082] In other words, in one example, during the process of adjusting the image features of the intelligent digital human based on the target trigger command indicated by the first intent feature, intelligent question answering can also be performed. For instance, in response to the first intent feature, an intelligent question-and-answer command can be generated, and based on this command, a target prompt word can be generated using the first intent feature and a preset prompt word template. Then, a large language model service can be invoked to obtain the first response information using the target prompt word. Furthermore, a text-to-speech service can be used to convert the first response information into audio data, which can then be used by the intelligent digital human for voice broadcasting.
[0083] Furthermore, in another example, before using the intelligent digital human for voice broadcasting, the mouth movements of the intelligent digital human can be rendered in real time based on the text data of the first response information, so that the mouth movements can match the text content and be displayed synchronously with the voice broadcast.
[0084] Furthermore, in one example, the target object can also respond to the content of the intelligent digital human's voice broadcast, and then generate a second response information based on the target object's response (i.e., the first response information), and then output the second response information, thus realizing intelligent question answering.
[0085] For example, in one instance, after the intelligent digital human is awakened, it can output initial response information. For instance, as shown in Figure 3(a), the target object triggers the intelligent digital human's image adjustment function, or after logging into the application (or platform, etc.) to adjust the intelligent digital human, the intelligent digital human outputs response information such as "self-introduction" and "guided content" to the target object, thereby guiding the target object to input control operations (such as inputting text, voice, etc. in a text box).
[0086] It should be noted that this disclosed solution can be adapted to various control operations (such as text, voice, or video), such as text input operations that require silence, or more intelligent voice interaction operations, thus meeting the needs of different users and improving the user experience.
[0087] Furthermore, in one example, the present disclosure provides other interactive functions during the dialogue between the target object and the intelligent digital human. For example, as shown in Figure 3(b), a function bar button (e.g., a "+" button) is set to expand the function bar. In addition to displaying the real-time dialogue function, the function bar also includes functions such as uploading pictures, sharing, and one-click adjustment. In this way, "communicating and interacting" with the intelligent digital human can be effectively realized, thereby enabling the intelligent digital human to quickly and accurately understand the user's needs and provide more timely and realistic feedback, thereby further improving the user experience.
[0088] Furthermore, in one example, to further enhance the user's sense of control, control buttons for intelligent question answering can be set. For example, as shown in Figure 3(c), the current conversation with the intelligent digital human can be started or ended in response to the target object's operation of the conversation start / stop button.
[0089] Furthermore, in one example, as shown in Figure 3(d), the present invention can also display the content of the conversation as subtitles in real time when the target object is having a conversation with the intelligent digital human; in addition, as shown in Figure 3(e), the present invention can also view the chat history between the target object and the intelligent digital human through the function of "expanding the input box", thereby improving the user experience.
[0090] In this way, the disclosed solution uses intelligent question-and-answer to adjust the intelligent digital human, thereby realizing "communication and interaction" between the target object and the intelligent digital human, which further enriches and enhances the user experience.
[0091] Figure 4 This is a schematic flowchart of an interaction method according to an embodiment of this application. This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices. It is understood that the above... Figures 1 to 3(e) The methods shown can also be applied to this example, but the related content will not be elaborated further in this example.
[0092] Furthermore, the method includes at least a portion of the following: For example... Figure 4 As shown, it includes:
[0093] Step S401: In response to the first control operation, obtain the first intent feature of the first control operation.
[0094] Here, the first intention feature is at least used to instruct the intelligent digital human to adjust its appearance features. For details on this part, please refer to the example above; it will not be repeated here.
[0095] Step S402: Based on the target trigger instruction indicated by the first intent feature, display at least two first product objects recommended by the first intent feature.
[0096] Step S403: In response to a second control operation, the second control operation is used to select a target product object from at least two first product objects.
[0097] Step S404: Using the target product object, adjust the image features of the intelligent digital human to display the intelligent digital human with adjusted image features.
[0098] Step S405: Output the first response information corresponding to the first intent feature.
[0099] It should be noted that after obtaining the first intent feature of the first control operation, a corresponding target trigger instruction can be executed according to the intent content indicated by the first intent feature. For example, if the intent content indicated by the first intent feature is used to indicate the recommendation of at least two first product objects (such as recommending a complete set of clothing), a target trigger instruction can be generated to indicate the recommended product objects. At this time, according to the target trigger instruction, at least two first product objects recommended can be displayed for the target object to select. Furthermore, in response to the second control operation (i.e., the selection operation) of the target object, the target product object is determined, and then the image features of the intelligent digital human are adjusted using the target product object, and the adjusted intelligent digital human is displayed.
[0100] For example, taking the first intent feature as an indication of recommended clothing, firstly, as shown in Figure 5(a), according to the recommendation instruction indicated by the first intent feature, multiple recommended clothing items are displayed on the display interface. For example, a display page containing multiple recommended clothing items (e.g., tops and bottoms) pops up in the right sidebar of the display interface; secondly, after the target object selects the top of style 1, the bottom of style 1, etc. from the displayed clothing, the object changes its clothing to obtain the smart digital human after changing clothes as shown in Figure 5(b); finally, the smart digital human after changing clothes can also be used to output the first response information.
[0101] For further explanation regarding the intelligent digital human's output of the first response information, please refer to the example above; it will not be repeated here.
[0102] It should be noted that in this example, after the target trigger command is executed, a prompt message indicating that the recommended preparation is complete can also be displayed, as shown in Figure 5(c), displaying the "Clothing Button". In this way, by clicking the Clothing Button, multiple recommended outfits can be displayed.
[0103] Furthermore, in one example, as shown in Figure 5(a), the present invention can also set up buttons such as clothing collection function, search function, and swipe function on the display page according to actual needs, so as to meet the different needs of users.
[0104] In this way, the disclosed solution can recommend products that meet the needs of the target audience based on intent characteristics, and then adjust the image characteristics of the intelligent digital human based on the target product selected by the target audience. This allows for a more intuitive demonstration of the effect of the target product on the intelligent digital human, enriching interactivity and ensuring that the target audience can have a comprehensive understanding of the product. This greatly enriches and enhances the user experience.
[0105] It should be noted that, in one example, the first intent feature can also directly instruct the intelligent digital human to change to a specified outfit. In this case, a change instruction can be directly generated, and the outfit to be changed can be determined based on the change instruction. Then, the image features of the intelligent digital human can be adjusted using the outfit, and the adjusted intelligent digital human can be displayed. This process can achieve the adjustment of image features without recommendation.
[0106] Alternatively, in another example, the first intention feature can also instruct the intelligent digital human to adjust its body movements. In this case, a body adjustment instruction can be generated based on the first intention feature, and the body movements to be displayed can be determined based on the body adjustment instruction to display the adjusted intelligent digital human.
[0107] It should be noted that the above is only a specific example of the first intent feature. In practical applications, there may be other intent features, and this disclosure does not limit them.
[0108] Furthermore, in a specific example, after the target object selects the target product object from among the displayed product objects, other product objects associated with the target product object can be further displayed, such as:
[0109] In response to the identification operation for the target product object, other second product objects associated with the description information of the target product object are displayed.
[0110] For example, in one example, the target user clicks the recognition button corresponding to the target product object on the display page to display other product objects (i.e., second product objects) associated with the description information of the target product object. For example, continuing with Figure 5(a), when adjusting the clothing of the smart digital human using the top of style 1 and displaying the adjusted smart digital human, as shown in Figure 5(d), the target user can also click the recognition button corresponding to the top of style 1. At this time, other styles of tops associated with the description information of the top of style 1 can be displayed in the style display area. In this way, the recommendation of similar product objects is realized, thereby further improving the user experience.
[0111] In this way, the disclosed solution utilizes the identification function of the target product object to obtain more information about the product object, thereby increasing the range of product objects that users can choose from, which helps users quickly find satisfactory product objects and further improves the user experience.
[0112] Furthermore, in a specific example, after the target object selects the target product object from the displayed multiple product objects, the method also includes displaying product feature information specific to the target product object. For example, it could specifically include:
[0113] In response to a product comparison operation for the target product object, at least two product feature information corresponding to the target product object are displayed.
[0114] Here, the product feature information includes at least one of the following: price, sales volume, etc. Furthermore, different product feature information comes from different recommendation sites (here, "site" can specifically refer to a website). For example, for the same product, the price of the same product on different recommendation sites may be different. In this case, the solution disclosed herein can intuitively display the product feature information of the same product on different recommendation sites, thereby further improving the user experience.
[0115] For example, continuing with Figure 5(d), after the target clicks the product comparison button (not shown) for the top of style 1, multiple product feature information of the top of style 1 will be displayed on multiple recommended sites, such as... Figure 6 As shown, the price (e.g., price 1) and sales volume (e.g., sales volume 1) of the top of style 1 on website 1 are displayed, as well as the price (e.g., price 2) and sales volume (e.g., sales volume 2) of the top of style 1 on website 2. In this way, the product feature information of the same product object under different recommended websites is compared and displayed, which makes it easier for users to make decisions quickly and accurately, thereby further improving the user experience.
[0116] In this way, the disclosed solution can utilize the comparison function of product feature information of product objects to display the product feature information of the same product object under different recommendation sites to the target object. This makes it easier for users to quickly find the product object they need based on the differences. Furthermore, by comparing the product feature information of product objects, it helps users make more accurate decisions, thereby further improving the user experience.
[0117] The following detailed explanation of the present disclosure solution is provided with reference to a specific example. In this example, as shown in Figures 7(a) and 7(b), the implementation steps of the present disclosure solution may specifically include:
[0118] Step S701: Obtain the control operation input by the target object, and input the obtained control operation into the large language model (that is, the second model mentioned above) to perform intent understanding and obtain the intent features of the target object (that is, the first intent features mentioned above).
[0119] In one example, the control operation can be a specific input operation, such as inputting a piece of text content. Alternatively, the control operation can be a voice operation, such as inputting control voice data through a microphone. In this case, a speech-to-text service can be invoked to generate a request instruction that understands the user's intent, thereby requesting the large language model to obtain the intent features of the target object. Further, the control voice data representing the control operation is converted into target text data, and the large language model service is invoked to input the target text data into the large language model to obtain the intent features of the target object.
[0120] Step S702: Generate the corresponding trigger command based on the intent characteristics of the target object.
[0121] For example, in one instance, if the intent feature of the target object is used to instruct the intelligent digital human (e.g., a 3D intelligent digital human) to change a specified outfit, a specific change instruction can be generated, and the process proceeds to step S703.
[0122] Alternatively, in another example, if the intent feature of the target object is used to indicate the recommendation of at least two outfits (e.g., recommending a complete outfit), then a recommendation instruction is generated and the process proceeds to step S704.
[0123] Alternatively, in another example, if the intent feature of the target object is used to instruct the intelligent digital human to adjust its limb movements, then a limb adjustment instruction is generated to trigger the intelligent digital human's limb movements, and the process proceeds to step S707.
[0124] It should be noted that in practical applications, different functions can be scheduled through central control, thereby realizing functions such as intelligent clothing recommendation, intelligent question answering, intelligent digital human dressing, and intelligent digital human driving limb and lip movements.
[0125] Step S703: Determine the target clothing to be replaced based on the replacement instruction indicated by the target object's intention characteristics; and proceed to step S706.
[0126] Step S704: Based on the recommendation instruction indicated by the intent characteristics of the target object, invoke the clothing recommendation service to display at least two recommended clothing items, thereby realizing intelligent clothing recommendation; and proceed to step S705.
[0127] Step S705: In response to the target object's selection operation on the displayed clothing, determine the target clothing to be changed; and proceed to step S706.
[0128] Step S706: Using the target clothing, adjust the image features of the intelligent digital human to obtain the adjusted intelligent digital human, and display the adjusted intelligent digital human. Proceed to step S708.
[0129] For example, the target clothing can be rendered onto a smart digital human to obtain a rendered smart digital human, which can then be displayed.
[0130] It should be noted that in this example, the intelligent digital human can also be displayed in 360 degrees. For example, while displaying the rendered intelligent digital human, it can also automatically rotate 360 degrees around a preset axis to display it in 360 degrees. This makes it easier to browse the effect of clothing changes from more dimensions, thereby effectively improving the user experience.
[0131] Alternatively, in one example, when displaying the rendered intelligent digital human, a 360-degree display instruction for the target object is obtained. At this time, based on the 360-degree display instruction, the intelligent digital human is triggered to automatically rotate 360 degrees around a preset axis.
[0132] Step S707: Based on the limb adjustment instructions indicated by the target object's intention characteristics, determine the target motion characteristics that the intelligent digital human needs to display. Then proceed to step S708.
[0133] Step S708: Generate intelligent question-and-answer instructions based on the intent characteristics of the target object; and proceed to step S709.
[0134] Step S709: Invoke the large language model service to infer the response content for responding to the target object based on the target object's intent characteristics. Proceed to step S710.
[0135] Step S710: Using a text-to-speech service, the response content for responding to the target object is converted into response speech data, and the intelligent digital human is used for speech broadcasting. At the same time, the lip movements of the intelligent digital human are rendered in real time based on the response speech data and displayed synchronously with the speech broadcast.
[0136] Step S711: Adjust the actions displayed by the intelligent digital human according to the target action features to be displayed (such as body movements and lip movements), and display the adjusted intelligent digital human in real time.
[0137] In this way, the disclosed solution can adjust the clothing or movements of the intelligent digital human according to the user's different intentions, and display the intelligent digital human after adjustment (such as changing clothes). At the same time, voice broadcast is used to respond to the user's response. Thus, the disclosed solution presents the user with different clothing matching effects through multi-dimensional interaction between the user and the intelligent digital human, and uses the intelligent digital human as a carrier, thereby enhancing the user's immersive experience.
[0138] This disclosure provides an interactive device, such as... Figure 8 As shown, it includes:
[0139] Processing unit 801 is configured to, in response to a first control operation, acquire a first intent feature of the first control operation, wherein the first intent feature is at least used to instruct the intelligent digital human to adjust its image features; adjust the image features of the intelligent digital human based on a target trigger instruction indicated by the first intent feature to obtain an adjusted intelligent digital human; and...
[0140] The output unit 802 is used to display the adjusted intelligent digital human and output first response information corresponding to the first intent feature.
[0141] In a specific example of the scheme disclosed herein, the processing unit is further configured to:
[0142] When displaying the adjusted intelligent digital human, trigger the adjusted intelligent digital human to rotate around a preset axis for multi-dimensional display;
[0143] Alternatively, in the case of displaying the adjusted intelligent digital human, in response to the display operation, the adjusted intelligent digital human can be triggered to rotate around a preset axis for multi-dimensional display.
[0144] In a specific example of the scheme disclosed herein, the processing unit is further configured to:
[0145] Based on the first intent feature of the first control operation and the preset prompt word template, a target prompt word is generated;
[0146] The target prompt word is input into the first model to obtain the first response information.
[0147] In a specific example of the scheme disclosed herein,
[0148] The processing unit is further configured to acquire first response information in response to the first response information; input the first response information into the first model at least to obtain second response information in response to the first response information; and
[0149] The output unit is also used to output the second response information for intelligent question answering.
[0150] In a specific example of the scheme disclosed herein,
[0151] The processing unit is specifically used to obtain at least two first product objects recommended by the first intent feature based on the target trigger instruction;
[0152] The output unit is specifically used to display at least two first product objects recommended by the first intent feature;
[0153] The processing unit is further specifically configured to respond to a second control operation, the second control operation being configured to select a target product object from at least two first product objects; and to adjust the image features of the intelligent digital human using the target product object to obtain an intelligent digital human with adjusted image features.
[0154] In a specific example of the scheme disclosed herein,
[0155] The processing unit is also configured to respond to an identification operation for the target product object;
[0156] The output unit is also used to display other second product objects associated with the description information of the target product object.
[0157] In a specific example of the scheme disclosed herein,
[0158] The processing unit is also configured to respond to a product comparison operation for the target product object;
[0159] The output unit is also used to display at least two product feature information corresponding to the target product object, wherein the different product feature information comes from different recommendation sites.
[0160] In a specific example of the disclosed solution, the processing unit is specifically used for:
[0161] In response to the first control operation, the operation features of the first control operation are input into the second model to obtain the first intention feature of the first control operation.
[0162] In a specific example of the scheme disclosed herein,
[0163] The processing unit is specifically configured to convert the first control operation into target text data in response to a first control operation containing voice data.
[0164] The target text data is input into the second model to obtain the first intent feature of the first control operation.
[0165] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.
[0166] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0167] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0168] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0169] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0170] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0171] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as interactive methods. For example, in some embodiments, the interactive method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the interactive method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform interactive methods by any other suitable means (e.g., by means of firmware).
[0172] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0173] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0174] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0176] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0177] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0178] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0179] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An interactive method applied to a fitting room scenario where an intelligent digital human can change clothes and display a three-dimensional dressing effect, the method comprising: In response to a first control operation, a first intent feature of the first control operation is acquired, wherein the first intent feature is at least used to instruct the intelligent digital human to adjust its image features and to instruct the intelligent digital human to adjust its limb movements. Based on the target triggering instruction indicated by the first intent feature, at least two first product objects are determined, so that the target object can select the target product object according to the at least two first product objects, and adjust the image features of the intelligent digital human based on the target product object to display the adjusted intelligent digital human; and based on the limb adjustment instruction indicated by the first intent feature, the limb movements that the intelligent digital human needs to display are determined, so as to adjust the movements displayed by the intelligent digital human according to the limb movements that need to be displayed. The intelligent digital human with adjusted image outputs first response information corresponding to the first intention feature to facilitate intelligent question answering; wherein, during the process of the first response information being broadcast and played by voice, the mouth movements of the intelligent digital human with adjusted image can match the text content corresponding to the first response information. The method further includes: When displaying the adjusted intelligent digital human, trigger the adjusted intelligent digital human to rotate 360 degrees around a preset axis for multi-dimensional display; Alternatively, in the case of displaying the adjusted intelligent digital human, in response to the display operation, the adjusted intelligent digital human can be triggered to rotate 360 degrees around a preset axis for multi-dimensional display.
2. The method according to claim 1, further comprising: Based on the first intent feature of the first control operation and the preset prompt word template, a target prompt word is generated; The target prompt word is input into the first model to obtain the first response information.
3. The method according to claim 2, further comprising: Obtain the first response information in response to the first response information; The first response information is input into the first model at least to obtain the second response information in response to the first response information; as well as Output the second response information for intelligent question answering.
4. The method according to any one of claims 1-3, wherein, The process of determining at least two first product objects based on the target trigger instruction indicated by the first intent feature, allowing the target object to select a target product object based on the at least two first product objects, and adjusting the image features of the intelligent digital human based on the target product object to display the adjusted intelligent digital human, includes: Based on the target trigger instruction, display at least two first product objects recommended by the first intent feature; In response to a second control operation, the second control operation is used to select a target product object from at least two first product objects; Using the target product object, adjust the image features of the intelligent digital human to display the intelligent digital human with adjusted image features.
5. The method according to claim 4, further comprising: In response to the identification operation for the target product object, other second product objects associated with the description information of the target product object are displayed.
6. The method according to claim 4, further comprising: In response to a product comparison operation for the target product object, at least two product feature information corresponding to the target product object are displayed, wherein the different product feature information comes from different recommendation sites.
7. The method according to any one of claims 1-3, wherein, The step of obtaining a first intent feature of the first control operation in response to the first control operation includes: In response to the first control operation, the operation features of the first control operation are input into the second model to obtain the first intention feature of the first control operation.
8. The method according to claim 7, wherein, In response to a first control operation, the operation features of the first control operation are input into a second model to obtain a first intent feature of the first control operation, including: In response to a first control operation containing voice data, the first control operation is converted into target text data; The target text data is input into the second model to obtain the first intent feature of the first control operation.
9. An interactive device, applied in a fitting room scenario where an intelligent digital human can change clothes and display a three-dimensional dressing effect, comprising: The processing unit is configured to, in response to a first control operation, acquire a first intent feature of the first control operation, wherein the first intent feature is at least used to instruct the intelligent digital human to adjust its image features and to instruct the intelligent digital human to adjust its limb movements; based on the target trigger instruction indicated by the first intent feature, determine at least two first product objects, so that a target object can select a target product object according to the at least two first product objects, and adjust the image features of the intelligent digital human based on the target product objects to obtain an adjusted intelligent digital human; and based on the limb adjustment instruction indicated by the first intent feature, determine the limb movements that the intelligent digital human needs to display, and adjust the movements displayed by the intelligent digital human according to the limb movements that need to be displayed. The output unit is used to display the adjusted intelligent digital human and to output the first response information corresponding to the first intention feature using the image-adjusted intelligent digital human, so as to facilitate intelligent question answering; wherein, during the process of the first response information being broadcast and played by voice, the mouth movements of the image-adjusted intelligent digital human can match the text content corresponding to the first response information. The processing unit is further configured to: When displaying the adjusted intelligent digital human, trigger the adjusted intelligent digital human to rotate 360 degrees around a preset axis for multi-dimensional display; Alternatively, in the case of displaying the adjusted intelligent digital human, in response to the display operation, the adjusted intelligent digital human can be triggered to rotate 360 degrees around a preset axis for multi-dimensional display.
10. The interactive device according to claim 9, wherein, The processing unit is further configured to: Based on the first intent feature of the first control operation and the preset prompt word template, a target prompt word is generated; The target prompt word is input into the first model to obtain the first response information.
11. The interactive device according to claim 10, wherein, The processing unit is further configured to acquire first response information in response to the first response information; input the first response information into the first model at least to obtain second response information in response to the first response information; and The output unit is also used to output the second response information for intelligent question answering.
12. The interactive device according to any one of claims 9-11, wherein, The processing unit is specifically used to obtain at least two first product objects recommended by the first intent feature based on the target trigger instruction; The output unit is specifically used to display at least two first product objects recommended by the first intent feature; The processing unit is further specifically configured to respond to a second control operation, the second control operation being configured to select a target product object from at least two first product objects; and to adjust the image features of the intelligent digital human using the target product object to obtain an intelligent digital human with adjusted image features.
13. The interactive device according to claim 12, wherein, The processing unit is also configured to respond to an identification operation for the target product object; The output unit is also used to display other second product objects associated with the description information of the target product object.
14. The interactive device according to claim 12, wherein, The processing unit is also configured to respond to a product comparison operation for the target product object; The output unit is also used to display at least two product feature information corresponding to the target product object, wherein the different product feature information comes from different recommendation sites.
15. The interactive device according to any one of claims 9-11, wherein, The processing unit is specifically used for: In response to the first control operation, the operation features of the first control operation are input into the second model to obtain the first intention feature of the first control operation.
16. The interactive device according to claim 15, wherein, The processing unit is specifically configured to convert the first control operation into target text data in response to a first control operation containing voice data. The target text data is input into the second model to obtain the first intent feature of the first control operation.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Virtual image generation method and device, electronic equipment and storage medium
CN114187394A
Method for providing commodity recommendation information and electronic equipment
CN116739699A
Information interaction method and device, electronic equipment and storage medium
CN118069798A