An AI real-time interaction method and system supporting user-defined digital human image

Through AI real-time interaction methods that support users to customize digital human images, the problem of digital human images lacking personalization and user-defined capabilities is solved, and the generation of user-defined digital human images and backgrounds is realized, user experience and interaction effects are improved, and technical application scenarios are expanded.

CN119179387BActive Publication Date: 2025-05-06BEIJING FIBERHOME TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411236410.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-05-06
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

In the existing AI real-time interaction technology, the digital human image lacks personalization and user customization capabilities, resulting in poor user experience and lack of flexibility in interaction methods, which limits the application expansion of the technology in more scenarios.

Method used

It provides an AI real-time interaction method that supports users to customize digital human images. By obtaining user custom information and user information, it generates digital human images and backgrounds that meet users' expectations, and captures user voice and action information in real time to conduct natural and smooth interactive responses.

Benefits of technology

It realizes that users customize the digital person image according to their preferences and needs, improves the personalization and intelligence of the overall layout of digital persons, improves the user experience and interaction effects, meets the users' growing personalized needs, and expands the application scenarios of AI real-time interaction technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119179387B_ABST
    Figure CN119179387B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of digital human interaction technology, and aims to solve the problem in the prior art that when a user interacts with a digital human, the digital human lacks personalization and user customization capabilities, and the interaction method lacks flexibility, resulting in poor user experience. The present invention provides an AI real-time interaction method and system that supports user customization of a digital human image, the method comprising: obtaining customized information and user information of the digital human image; determining the digital human image based on the customized information of the digital human image; generating a digital human background based on the user information and the determined digital human image; obtaining the voice interaction information input by the user in real time and capturing the user's action information; interacting with the user in real time based on the voice interaction information input by the user; and determining the digital human's response to the user's interaction action based on the user's action information. The present invention can customize the digital human image according to the user's preferences, and the digital human can make smooth and natural responses during interaction, thereby improving the user's interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human interaction technology, and in particular to an AI real-time interaction method and system that supports user-defined digital human images. Background Art

[0002] With the rapid development and popularization of artificial intelligence technology, AI real-time interaction technology has become an important application in many fields, such as virtual customer service, online education, entertainment interaction, and virtual reality. This technology provides users with a more convenient and efficient service experience by simulating human interaction behavior, greatly enriching the interactivity and immersion of the digital world. However, despite the significant progress made in AI real-time interaction technology, traditional interaction methods still face many challenges, the most prominent of which is the lack of personalization and user customization capabilities.

[0003] Although the existing digital human images have achieved a certain degree of diversity in appearance, behavior patterns and interaction methods, most of them are still preset by the system. Users can only passively accept them during the interaction process and cannot customize them according to their preferences and needs. This lack of flexibility and personalized interaction not only limits the improvement of user experience, but also hinders the application and expansion of AI real-time interaction technology in more scenarios. In modern society, personalized needs are becoming increasingly prominent, and users have higher and higher expectations for digital human images that can reflect their unique tastes and styles. They hope that digital humans can not only understand and respond to their voices and movements, but also resonate with them in appearance, personality, behavior patterns, etc., so as to achieve a more in-depth and real interactive experience. For example, in the field of education, students may prefer to interact with a digital human teacher who is friendly and patient; in the field of entertainment, users may prefer to play games or chat with a digital human character with a specific style or appearance.

[0004] In view of this, the art needs an AI real-time interaction method and system that supports user-defined digital human images to solve the above problems. Summary of the invention

[0005] In order to solve the above technical problems, that is, to solve the problem in the prior art that when a user interacts with a digital human, the digital human lacks personalization and user customization capabilities, and the interaction method lacks flexibility, resulting in poor user experience.

[0006] In one aspect, the present invention provides an AI real-time interaction method that supports user-defined digital human images, the method comprising:

[0007] S1: Obtaining the custom information and user information of the digital human image;

[0008] S2: Determine the digital human image based on the customized information of the digital human image;

[0009] S3: Generate a digital human background based on user information and the determined digital human image;

[0010] S4: Acquire the voice interaction information input by the user and capture the user's action information in real time;

[0011] S5: interacting with the user in real time based on the voice interaction information input by the user;

[0012] S6: Determine the interactive action of the digital human in response to the user based on the action information of the user.

[0013] In some preferred implementations, the user information includes the user's identity information, the user's current scenario information, and the user's historical preference information; step S3 specifically includes:

[0014] S31: Determine whether the user's historical preference information is pre-stored in the system;

[0015] S321: If the user's historical preference information is pre-stored in the system, the user's identity information, the user's current scene information, the user's historical preference information and the determined digital human image are input into a first digital human background generation model, and a digital human background is generated based on the first digital human background generation model;

[0016] S322: If the user's historical preference information is not pre-stored in the system, the user's identity information, the user's current scene information and the determined digital human image are input into a second digital human background generation model, and a digital human background is generated based on the second digital human background generation model.

[0017] In some preferred embodiments, it is characterized in that the first digital human background generation model is:

[0018]

[0019] Among them, B1 is the multidimensional vector of the output digital human background, the number of dimensions is b1, K1 is the number of items to be summed, J1 is the number of layers, ω 1k ,θ 1jk , α 1ij , β 1ij , γ 1ij and δ 1ij are weights, ∈ 1j and ∈0 are bias terms, d1 is the determined digital human image vector The number of dimensions, D i is the i-th dimension of the digital human image vector D, i1 is the user's identity information vector The number of dimensions, Ii is the user's identity information vector The i-th dimension of , s1 is the user's current scene information vector The number of dimensions, S i is the user's current scene information vector The i-th dimension of , h1 is the number of dimensions of the user's historical preference information vector H, H i is the user's historical preference information vector The i-th dimension of .

[0020] In some preferred embodiments, it is characterized in that the second digital human background generation model is:

[0021]

[0022] Among them, B2 is the multi-dimensional vector of the output digital human background, the number of dimensions is b2, L1 is the number of convolution operations, M1 is the number of transformer operations, is the digital human image vector, is the user's identity information vector, is the user's current scene information vector, η 1m is the weight, and ∈1 is the bias term.

[0023] In some preferred embodiments, step S5 specifically includes:

[0024] S51: Converting the voice interaction information input by the user into text information;

[0025] S52: performing natural language processing on the converted text information and extracting key information;

[0026] S53: Retrieve relevant information from a local knowledge base and / or a cloud database according to the extracted key information;

[0027] S54: generating response content according to the retrieved relevant information and the determined digital human image;

[0028] S55: converting the generated response content into a voice signal and giving a voice response to the user;

[0029] S56: Acquire the reaction information of the user after receiving the voice response, and adjust the subsequent real-time interaction according to the user's reaction information.

[0030] In some preferred embodiments, in step S52, the key information is extracted using the following formula:

[0031]

[0032] Among them, Ξ is the extracted key information set, N is the number of words in the text information, is the weight of the ιth word, Θ ι (Ψ) is the result of the basic natural language processing function operation on the ιth word, M is the number of first-class factors affecting the word weight, For the The weight coefficient of the first type of factors is For the ιth word in The characteristic function results under the first-class factors, P is the number of second-class factors affecting the word weight, μ χ is the weight coefficient of the xth second-category factor, Γ χ (Ψ ι ) is the characteristic function result of the ιth word under the χth second-category factor, and Ψ is the converted text information;

[0033] The first category of factors includes at least two of the importance of the word in a specific subject area, the relevance of the word to the current interaction scenario, the association of the word with the current hot topic, the ambiguity of the word in different contexts, and the fit of the word with the user's specific interest area. The second category of factors includes at least two of the user's historical interaction preferences, the current time, the user's geographic location, and the user's language habits.

[0034] In some preferred implementations, in step S54, the response content is generated using the following formula:

[0035]

[0036] in, is the generated response content, α′ is the weight of the digital human image feature, Δ is the digital human image feature vector, β′ is the weight of the retrieved information feature, Σ is the feature vector of the relevant information retrieved from the local knowledge base and / or the cloud database, γ′ is the weight of the context information feature, Ω is the context information feature vector of the current interaction, Q is the number of first-class interaction adjustment factors that affect the response content, and η λ is the weight of the λth first-class interaction adjustment factor, Λ λ (Δ, Σ, Ω) is the result of the λth first-class interactive adjustment function under the joint action of the digital human image feature vector, the retrieval information feature vector and the context information feature vector, S is the number of second-class interactive adjustment factors that affect the response content, ξ ρ is the weight of the ρth second-type interaction adjustment factor, Ψ ρ (Δ, Σ, Ω) is the result of the ρth second-class interactive adjustment function under the joint action of the digital human image feature vector, the retrieval information feature vector and the context information feature vector;

[0037] The first type of interaction adjustment factors includes at least two of the user's emotional state, the urgency of the interaction, the strength of the user's emotional tendency, the clarity of the purpose of the interaction, and the current task status of the digital human. The second type of interaction adjustment factors includes at least two of the external environmental factors, popular trends, social and cultural trends, industry dynamics, and the growth stage of the digital human.

[0038] In some preferred embodiments, step S6 specifically includes:

[0039] S61: parsing the action information of the user, wherein the action information includes body movement information and facial expression change information;

[0040] S62: judging the user's action intention according to the parsed action information;

[0041] S63: determining the interactive action of the digital human in response to the user according to the action intention of the user;

[0042] S64: Optimizing the digital human's response to the user's interactive actions.

[0043] In some preferred implementations, in step S64, the following formula is used to optimize the digital human's response to the user's interactive action:

[0044]

[0045] Among them, D optimized For the optimized digital human interaction action, D action_i is the user’s ith sub-action, n is the total number of sub-actions, m is the number of frames in the time series, and t j is the time point of the jth frame, D action (t j ) is the time t j Digital human action, D action (t j-1 ) is the time t j-1 Digital human action, D max_change is the preset maximum action change range, k is the total number of actions in the real human action database, H action_i is the ith action in the real human action database, Sim is the similarity calculation function, C context_detail is the detailed environment information of the interaction context, D characteristics_detail is the detailed characteristic parameter of the digital human, A constraints is the constraint condition of the action, D action_history The action history of the digital human, U expectation The user's desired action.

[0046] In another aspect, the present invention further provides an AI real-time interactive system that supports user-defined digital human images, the system comprising:

[0047] A first information acquisition module, which is used to acquire the custom information and user information of the digital human image;

[0048] A digital human image determination module, which is used to determine the digital human image based on the custom information of the digital human image;

[0049] A digital human background generation module, which is used to generate a digital human background based on user information and a determined digital human image;

[0050] The second information acquisition module is used to acquire the voice interaction information input by the user and capture the user's action information in real time;

[0051] A real-time interaction module, which is used to interact with the user in real time based on the voice interaction information input by the user;

[0052] A digital human action determination module is used to determine the interactive action of the digital human in response to the user based on the action information of the user.

[0053] As can be seen from the above, the AI ​​real-time interaction method and system that supports user-defined digital human images provided by the present invention has the following beneficial technical effects:

[0054] The present invention enables users to freely customize the image of the digital human according to their own preferences and needs, thereby creating a digital human image that meets personal expectations and preferences. The background of the digital human can be automatically generated based on user information and the digital human image, thereby improving the personalization and intelligence of the overall layout of the digital human. At the same time, the present invention also needs to have the ability to capture user voice and action information in real time, so that the digital human can make a more natural and smooth interactive response based on the user's real-time input, improve flexibility, further enhance user experience and interactive effects, meet users' growing personalized needs, and expand the application scenarios of AI real-time interactive technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0056] Figure 1 A flowchart of the AI ​​real-time interaction method of the present invention supporting user-defined digital human images;

[0057] Figure 2 The figure is a schematic diagram of the structure of the AI ​​real-time interactive system supporting user-defined digital human images of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0059] Based on the problem pointed out in the background technology that the digital human in the prior art lacks the ability of personalization and user customization when the user interacts with the digital human, and the interaction mode lacks flexibility, resulting in poor user experience. The present invention provides an AI real-time interaction method and system that supports user customization of the digital human image, aiming to allow users to freely customize the image of the digital human according to their preferences and needs, thereby creating a digital human image that meets personal expectations and preferences, and improving the personalization of the overall layout of the digital human. At the same time, when the user interacts with the digital human, the digital human can make a more natural and smooth interactive response, further improving the user experience and interaction effect.

[0060] like Figure 1 As shown, the AI ​​real-time interaction method of the present invention supporting user-defined digital human images includes:

[0061] S1: Obtaining the custom information and user information of the digital human image;

[0062] S2: Determine the digital human image based on the customized information of the digital human image;

[0063] S3: Generate a digital human background based on user information and the determined digital human image;

[0064] S4: Acquire the voice interaction information input by the user and capture the user's action information in real time;

[0065] S5: interacting with the user in real time based on the voice interaction information input by the user;

[0066] S6: Determine the interactive action of the digital human in response to the user based on the user's action information.

[0067] In the above, the customized information of the digital human image may include the digital human's appearance (such as hairstyle, clothing, facial expression, etc.), personality and behavior pattern (such as friendliness, humor, rigor, etc.), and may also include the digital human's voice, gender, age and emotional expression, etc., which can be flexibly set by those skilled in the art in practical applications. In addition, the acquired user information may include the user's identity information, the user's current scene information and the user's historical preference information, and may also include the user's language information and interest information, etc., which can be flexibly set by those skilled in the art.

[0068] In a preferred embodiment, the user information includes the user's identity information, the user's current scene information and the user's historical preference information. The above step S3 specifically includes:

[0069] S31: Determine whether the user's historical preference information is pre-stored in the system;

[0070] S321: If the user's historical preference information is pre-stored in the system, the user's identity information, the user's current scene information, the user's historical preference information and the determined digital human image are input into the first digital human background generation model, and the digital human background is generated based on the first digital human background generation model;

[0071] S322: If the user's historical preference information is not pre-stored in the system, the user's identity information, the user's current scene information and the determined digital human image are input into the second digital human background generation model, and the digital human background is generated based on the second digital human background generation model.

[0072] In actual applications, the system can retrieve the user's historical preference information, which can be constructed based on the user's previous interaction process or obtained through big data. In this case, the digital human background can be generated through the above-mentioned step S321. Of course, in the case where the user's historical preference information is not retrieved, the digital human background can be generated through the above-mentioned step S322. Through such a setting, that is, by considering the user's historical preference information, the system can generate a digital human background that is more in line with the user's personalized needs, thereby improving the user's interactive experience. At the same time, regardless of whether the user's historical preference information is pre-stored in the system, it can respond flexibly, ensuring the adaptability and stability of the system in different situations. In addition, the system determines which background generation model to use by judging whether the user's historical preference information exists. This design can efficiently utilize user information and avoid unnecessary calculations and resource waste. Since the system can generate the digital human background in real time based on user information and scene information, this greatly enhances the real-time interaction ability of the digital human and ensures the continuity and real-time nature of the interaction. Moreover, by providing a more personalized and user-friendly digital human background, it helps to improve user satisfaction and trust.

[0073] Further preferably, the first digital human background generation model is:

[0074]

[0075] Among them, B1 is the multidimensional vector of the output digital human background, the number of dimensions is b1, K1 is the number of items to be summed, J1 is the number of layers, ω 1k ,θ 1jk , α 1ij , β1ij , γ 1ij and δ 1ij are weights, ∈ 1j and ∈0 are bias terms, d1 is the determined digital human image vector The number of dimensions, D i Digital human image vector The i-th dimension of , i1 is the user's identity information vector The number of dimensions, I i is the i-th dimension of the user's identity information vector I, and s1 is the user's current scene information vector The number of dimensions, S i is the user's current scene information vector The i-th dimension of h1 is the user's historical preference information vector The number of dimensions, H i is the user's historical preference information vector The i-th dimension of ;

[0076] The calculation process of the above formula is: for each j, calculate the internal linear combination and activation function, is the linear unit activation function, and then for each k, the linear combination and tanh activation function operation are performed again, which is the hyperbolic tangent activation function. Finally, all k items are summed up, and the bias term ∈ 0 is added to obtain the final multi-dimensional vector B1 of the digital human background. Among them, α 1ij , β 1ij , γ 1ij and δ 1ij are used to adjust the weights of the various dimensions of the digital human image, user identity information, user current scene information, and user historical preference information in the linear combination of different levels, θ 1jk Used to adjust the weights of the tanh activation function outputs at different levels, ω 1k Used to adjust the weights of different items in the final output, ∈ 1j and ∈ 0 are used to adjust the offsets of different levels and the final output.

[0077] Optionally, each dimension of the output digital human background vector B1 represents different elements of the digital background, such as the RGB value of the color, the size of the background object, the position coordinates, texture features, etc.; these dimensional values ​​can be specifically mapped and converted and applied to the generation of the digital human background; for example, the values ​​of certain dimensions can be directly used as color parameters, and the color of the background can be determined by converting the RGB color space; the values ​​of other dimensions can be mapped to the position coordinate space according to certain rules to determine the position of the object in the background; the texture features can generate corresponding texture patterns through specific algorithms and applied to the background.

[0078] Further preferably, the second digital human background generation model is:

[0079]

[0080] Among them, B2 is the multi-dimensional vector of the output digital human background, the number of dimensions is b2, L1 is the number of convolution operations, M1 is the number of transformer operations, is the digital human image vector, is the user's identity information vector, is the user's current scene information vector, η 1m is the weight, ∈1 is the bias term;

[0081] The calculation process of the above formula is:

[0082] First, perform a convolution block operation (ConvBlock). For each l, perform a convolution block operation The digital human image vector, user identity information vector and user current scene information vector are integrated and feature extracted: the three vectors are spliced ​​together: Perform the first convolution operation Conv1 l =Conv2D(σ1(X l )), where Conv2D is a two-dimensional convolution operation, σ1 is an activation function (such as LeakyReLU); perform pooling operation: Pool1 l =MaxPooling2D(Conv1 l ), perform the maximum pooling operation to reduce the feature dimension; perform the second layer of convolution operation: Conv2 l =Conv2D(σ2(Pool1 l )), σ2 is the activation function, and convolution and activation function operations are performed again to extract more advanced features; the above process is repeated many times to obtain the final convolution block output;

[0083] Secondly, the attention mechanism sums the outputs of all convolutional blocks: The weighted processing is performed through the attention mechanism function Attention. The calculation of the attention mechanism is assumed to be Attention(Z)=softmax(W Z Z)·V Z Z, where W Z and V Z is a parameter matrix, and the softmax function converts the output into a probability distribution so that different features get different weights;

[0084] Again, the transformer block operation (TransformerBlock), for each m, the transformer block operation is performed Further capture the long-distance dependencies between input vectors: MultiHeadAttention(Q0,K0,V0)=Concat(head1,...,head h )W O ,in, Q0, K0 and V0 are query, key and value matrices respectively, and W O is the parameter matrix, h is the number of heads; feedforward neural network: FFN(x)=max(0,xW1+c1)W2+c2, W1, W2, c1 and c2 are parameters; layer normalization: Where E[x] is the mean, Var[x] is the variance, ∈ is a constant, and γ and β are parameters;

[0085] Finally, the output of the attention mechanism is summed with the output of the transformer block: Then add the bias term ∈1 to get the final output vector B2, η 1m It is used to adjust the weight of the transformer block output in the final output, and ∈1 is used to adjust the offset of the final output.

[0086] Optionally, similar to the aforementioned B1, each dimension of the output digital human background vector B2 also represents different elements of the digital background; the values ​​of these dimensions can be appropriately processed and converted according to the specific application scenario to generate the background picture in which the digital human is located; for example, the values ​​of certain dimensions can be used as shape parameters of the background object, and the corresponding shapes can be generated through a specific graphics generation algorithm; the values ​​of other dimensions can control the illumination intensity and direction of the background, and adjust the illumination effect of the background through the illumination model; different background textures and patterns can also be selected according to the values ​​of different dimensions to increase the richness and realism of the background.

[0087] In a preferred embodiment, the aforementioned step S5 specifically includes:

[0088] S51: Converting the voice interaction information input by the user into text information;

[0089] S52: performing natural language processing on the converted text information and extracting key information;

[0090] S53: Retrieve relevant information from a local knowledge base and / or a cloud database according to the extracted key information;

[0091] S54: generating response content according to the retrieved relevant information and the determined digital human image;

[0092] S55: converting the generated response content into a voice signal and giving a voice response to the user;

[0093] S56: Acquire the reaction information of the user after receiving the voice response, and adjust the subsequent real-time interaction according to the user's reaction information.

[0094] Through such settings, that is, by defining the voice interaction process in detail, efficient processing and intelligent response to user input voice is achieved. Specifically, it converts voice information into text, extracts key information using natural language processing technology, and then retrieves relevant information from local or cloud databases to generate and output targeted voice responses. In addition, by obtaining and analyzing user response information, dynamic adjustment of subsequent interaction content is achieved, thereby significantly improving the intelligence and personalization of the interactive experience.

[0095] Preferably, in the above step S52, the key information is extracted using the following formula:

[0096]

[0097] Among them, Ξ is the extracted key information set, N is the number of words in the text information, is the weight of the ιth word, Θ ι (Ψ) is the result of the basic natural language processing function operation on the ιth word, M is the number of first-class factors affecting the word weight, For the The weight coefficient of the first type of factors is For the ιth word in The characteristic function results under the first-class factors, P is the number of second-class factors affecting the word weight, μ χ is the weight coefficient of the xth second-category factor, Γ χ (Ψ ι) is the characteristic function result of the ιth word under the χth second-category factor, and Ψ is the converted text information; the first-category factors include at least two of the importance of the word in a specific subject area, the relevance of the word to the current interaction scene, the relevance of the word to the current hot topic, the polysemy of the word in different contexts, and the fit between the word and the user's specific interest area; the second-category factors include at least two of the user's historical interaction preferences, the current time, the user's geographical location, and the user's language habits. In a possible implementation, the text information converted from the voice interaction information input by the user is "Today's weather is very good, suitable for a walk in the park". After natural language processing, the three words "weather", "park", and "walk" are given higher weights and extracted as key information through function operations; in the subsequent step S53, the current weather condition information can be retrieved according to "weather", and the information of nearby parks suitable for walking can be retrieved according to "park" and "walk", etc., to prepare for the subsequent generation of response content. In another possible implementation, the text information converted from the voice interaction information input by the user is "A new smart device with a very high-tech feel has been developed today." In this embodiment, the words "research and development," "smart device," and "sense of technology" are given higher weights in the interactive scenarios of technology themes; these key information are finally extracted by calculating the characteristic function results and weight coefficients under various factors; in the subsequent step S53, relevant information such as introductions to the latest smart devices, technology development trends, etc. can be retrieved from the knowledge base and database based on the key information of "research and development," "smart device," and "sense of technology."

[0098] Preferably, in the above step S54, the response content is generated using the following formula:

[0099]

[0100] in, is the generated response content, α′ is the weight of the digital human image feature, Δ is the digital human image feature vector, β′ is the weight of the retrieved information feature, Σ is the feature vector of the relevant information retrieved from the local knowledge base and / or the cloud database, γ′ is the weight of the context information feature, Ω is the context information feature vector of the current interaction, Q is the number of first-class interaction adjustment factors that affect the response content, and η λ is the weight of the λth first-class interaction adjustment factor, Λ λ (Δ, Σ, Ω) is the result of the λth first-class interactive adjustment function under the joint action of the digital human image feature vector, the retrieval information feature vector and the context information feature vector, S is the number of second-class interactive adjustment factors that affect the response content, ξ ρ is the weight of the ρth second-type interaction adjustment factor, Ψ ρ(Δ, Σ, Ω) is the result of the ρth second-class interaction adjustment function under the joint action of the digital human image feature vector, the retrieval information feature vector and the context information feature vector; the first-class interaction adjustment factors include at least two of the user's emotional state, the urgency of the interaction, the strength of the user's emotional tendency, the clarity of the purpose of the interaction and the current task state of the digital human, and the second-class interaction adjustment factors include at least two of the external environmental factors, popular trends, social and cultural trends, industry dynamic changes and the growth stage of the digital human. In a possible implementation, the determined digital human image is a gentle and lovely girl, the relevant information retrieved is an introduction to a certain tourist attraction, and the current interaction context is that the user is planning a trip; the digital human image feature vector Δ contains the numerical representation of the gentle and lovely features; the retrieval information feature vector Σ contains the numerical representation of the name and characteristic information of the attraction; the context information feature vector Ω contains the numerical representation of the user's interest in travel and other information; by adjusting the weight coefficients α′, β′, γ′, the response content is generated, for example: "Dear! That attraction is very good. I heard that the scenery there is very beautiful and very suitable for people like you who love traveling!" Then in the subsequent step S55, this response content is converted into a voice signal and played to the user. In another possible implementation, the determined digital human image is a lively, cute, and humorous teenager, the retrieved relevant information is an introduction to a popular movie, and the current interactive context is that the user discusses entertainment activities with friends on weekends; the digital human image feature vector Δ contains numerical representations of lively and humorous features; the retrieval information feature vector Σ contains numerical representations of the name, type, and evaluation information of the movie; the context information feature vector Ω contains numerical representations of weekends, friends, and entertainment; by adjusting the basic weight coefficients α′, β′, γ′ and considering various interactive adjustment factors, response content is generated, for example: "Hey! It must be great to watch that popular movie with friends on the weekend! I heard that the plot of the movie is particularly exciting and full of surprises!" Then in the subsequent step S55, this response content is converted into a voice signal and played to the user.

[0101] In a preferred embodiment, the aforementioned step S6 specifically includes:

[0102] S61: parsing the user's action information, wherein the action information includes body movement information and facial expression change information;

[0103] S62: judging the user's action intention according to the parsed action information;

[0104] S63: determining the interactive action of the digital human in response to the user according to the action intention of the user;

[0105] S64: Optimize the digital human's response to the user's interactive actions.

[0106] Through such settings, the action interaction experience between digital humans and users can be improved. Specifically, by finely analyzing the user's body movements and facial expression changes, the user's action intentions can be accurately judged, and the digital human's response actions can be determined accordingly. In addition, the digital human's interactive actions are optimized to make the digital human's movements more natural and smooth, and the action interaction between the digital human and the user more tacit and vivid. This greatly enhances the interactivity and immersion between the digital human and the user, and improves the quality of the overall interactive experience.

[0107] In the above, it is preferred to respond to both the user's body movement information and facial expression change information. For example, when the user makes a waving action, the digital human can make a waving response action accordingly. The amplitude and speed of the action can be adjusted according to the intensity of the user's action to maintain the naturalness of the interaction. When the user makes an indication action, the digital human can turn its head or body along the indication direction to show attention and response; for example, according to the user's emotional state, the digital human adjusts its facial expression. When the user shows happiness, the digital human can smile; when the user looks confused, the digital human can frown and look thoughtful. In addition, the digital human's facial expression changes can be more delicate, such as slightly raising the corners of the mouth, blinking, etc., to enhance the affinity of the interaction; for example, when the user turns his head, the digital human can follow the direction of the user's head movement and maintain eye contact. When the user nods, the digital human can also gently nod in response to show understanding and agreement.

[0108] Preferably, in the above step S64, the following formula is used to optimize the digital human's response to the user's interactive action:

[0109]

[0110] Among them, D optimized For the optimized digital human interaction action, D action_i is the user’s ith sub-action, n is the total number of sub-actions, m is the number of frames in the time series, and t j is the time point of the jth frame, D action (t j ) is the time t j Digital human action, D action (t j-1 ) is the time t j-1 Digital human action, D max_change is the preset maximum action change range, k is the total number of actions in the real human action database, H action_i is the ith action in the real human action database, Sim is the similarity calculation function, C context_detail is the detailed environment information of the interaction context, D characteristics_detailis the detailed characteristic parameter of the digital human, A constraints is the constraint condition of the action, D action_history The action history of the digital human, U expectation is the user's desired action. Used to calculate the smoothness of the action, the value range is between 0 and 1, 1 means completely smooth, Used to calculate the realism of an action.

[0111] It should be noted that the method of the present invention can be performed by a single device, such as a computer or a server. The method of the present invention can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the present invention, and the multiple devices will interact with each other to complete the described method.

[0112] It should be noted that the above specific embodiments of the present invention are described. Other embodiments are within the scope of the present invention. In some cases, the actions or steps recorded in the present invention can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0113] Based on the same purpose, corresponding to any of the above-mentioned embodiments and methods, the embodiment of the present invention also provides an AI real-time interaction method that supports user-defined digital human images, such as Figure 2 As shown, the system includes:

[0114] A first information acquisition module 100, which is used to acquire the custom information and user information of the digital human image;

[0115] A digital human image determination module 200, which is used to determine the digital human image based on the custom information of the digital human image;

[0116] A digital human background generation module 300, which is used to generate a digital human background based on user information and a determined digital human image;

[0117] The second information acquisition module 400 is used to acquire the voice interaction information input by the user and capture the user's action information in real time;

[0118] A real-time interaction module 500, which is used to interact with the user in real time based on the voice interaction information input by the user;

[0119] The digital human action determination module 600 is used to determine the interactive action of the digital human in response to the user based on the user's action information.

[0120] The system of the above embodiment is used to implement the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0121] For the convenience of description, the above system is described by dividing it into various units according to its functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0122] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, systems, etc. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0123] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0124] Each embodiment of the present invention is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0125] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention is limited to these examples. Under the concept of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0126] While the invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications and variations of these embodiments will be apparent to those skilled in the art in light of the foregoing description.

[0127] One or more embodiments of the present invention are intended to cover all such replacements, modifications and variations that fall within the broad scope of the present invention. Therefore, any omissions, modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of the present invention should be included in the protection scope of the present invention.

Claims

1. An AI real-time interaction method supporting user-defined digital human images, characterized in that: The method comprises: S1: Obtaining the customized information and user information of the digital human image, wherein the user information includes the user's identity information, the user's current scene information, and the user's historical preference information; S2: Determine the digital human image based on the customized information of the digital human image; S3: Generate a digital human background based on user information and the determined digital human image; Step S3 specifically includes: S31: Determine whether the user's historical preference information is pre-stored in the system; S321: If the user's historical preference information is pre-stored in the system, the user's identity information, the user's current scene information, the user's historical preference information and the determined digital human image are input into a first digital human background generation model, and a digital human background is generated based on the first digital human background generation model; The first digital human background generation model is: ; in, is the multidimensional vector of the output digital human background, the number of dimensions is , is the number of terms to be summed, is the number of layers, , , , , and are weights, and are all bias terms, To determine the digital human image vector The number of dimensions, Digital human image vector No. dimensions, is the user's identity information vector The number of dimensions, is the user's identity information vector No. dimensions, is the user's current scene information vector The number of dimensions, is the user's current scene information vector No. dimensions, is the user's historical preference information vector The number of dimensions, is the user's historical preference information vector No. Dimensions; S322: If the user's historical preference information is not pre-stored in the system, the user's identity information, the user's current scene information and the determined digital human image are input into a second digital human background generation model, and a digital human background is generated based on the second digital human background generation model; The second digital human background generation model is: ; in, is the multidimensional vector of the output digital human background, the number of dimensions is , is the number of convolution operations, is the number of transformer operations, is the digital human image vector, is the user's identity information vector, is the user’s current scene information vector, is the weight, is the bias term; S4: Acquire the voice interaction information input by the user and capture the user's action information in real time; S5: interacting with the user in real time based on the voice interaction information input by the user; S6: Determine the interactive action of the digital human in response to the user based on the action information of the user.

2. The AI ​​real-time interaction method supporting user-defined digital human images according to claim 1, characterized in that: Step S5 specifically includes: S51: Converting the voice interaction information input by the user into text information; S52: performing natural language processing on the converted text information and extracting key information; S53: Retrieve relevant information from a local knowledge base and / or a cloud database according to the extracted key information; S54: generating response content according to the retrieved relevant information and the determined digital human image; S55: converting the generated response content into a voice signal and giving a voice response to the user; S56: Acquire the reaction information of the user after receiving the voice response, and adjust the subsequent real-time interaction according to the user's reaction information.

3. The AI ​​real-time interaction method supporting user-defined digital human images according to claim 2, characterized in that: In step S52, the key information is extracted using the following formula: ; in, is the key information set extracted, is the number of words in the text information, For the The weight of a word, For the The result of performing basic natural language processing functions on words. is the number of first-class factors that affect word weights, For the The weight coefficient of the first type of factors is For the The word in The characteristic function results under the first type of factors are: is the number of the second type of factors that affect the word weight, For the The weight coefficient of the second type of factors is For the The word in The characteristic function results under the second type of factors are: is the converted text information; The first category of factors includes at least two of the importance of the word in a specific subject area, the relevance of the word to the current interaction scenario, the association of the word with the current hot topic, the ambiguity of the word in different contexts, and the fit of the word with the user's specific interest area. The second category of factors includes at least two of the user's historical interaction preferences, the current time, the user's geographic location, and the user's language habits.

4. The AI ​​real-time interaction method supporting user-defined digital human images according to claim 2, characterized in that: In step S54, the response content is generated using the following formula: ; in, The generated response content is is the weight of the digital human image features, is the digital human image feature vector, is the weight of the retrieval information feature, is the feature vector of relevant information retrieved from the local knowledge base and / or cloud database, is the weight of the context information feature, is the context information feature vector of the current interaction, The number of first-order interaction adjustment factors that affect the content of the response, For the The weights of the first-order interaction adjustment factors, is the first The first kind of interactive adjustment function results, The number of second-type interaction adjustment factors that affect the content of the response, For the The weights of the second type of interaction adjustment factors, is the first The second type of interactive adjustment function results; The first type of interaction adjustment factors includes at least two of the user's emotional state, the urgency of the interaction, the strength of the user's emotional tendency, the clarity of the purpose of the interaction, and the current task status of the digital human. The second type of interaction adjustment factors includes at least two of the external environmental factors, popular trends, social and cultural trends, industry dynamics, and the growth stage of the digital human.

5. The AI ​​real-time interaction method supporting user-defined digital human images according to claim 1, characterized in that: Step S6 specifically includes: S61: parsing the action information of the user, wherein the action information includes body movement information and facial expression change information; S62: judging the user's action intention according to the parsed action information; S63: determining the interactive action of the digital human in response to the user according to the action intention of the user; S64: Optimizing the digital human's response to the user's interactive actions.

6. The AI ​​real-time interaction method supporting user-defined digital human images according to claim 5, characterized in that: In step S64, the following formula is used to optimize the digital human's response to the user's interactive action: ; in, For the optimized digital human interaction actions, For the user's Sub-actions, is the total number of sub-actions, is the number of frames in the time series, For the The time point of the frame, For in time Digital human action, For in time Digital human action, is the preset maximum action change range, is the total number of actions in the real human action database, The first Actions, is the similarity calculation function, Detailed environment information for the interaction context, are the detailed characteristic parameters of the digital human, are the constraints of the action, It is the action history record of the digital human. The user's desired action.

7. An AI real-time interactive system that supports user-defined digital human images, characterized in that: The system comprises: A first information acquisition module, which is used to acquire the customized information and user information of the digital human image, wherein the user information includes the user's identity information, the user's current scene information and the user's historical preference information; A digital human image determination module, which is used to determine the digital human image based on the custom information of the digital human image; A digital human background generation module, which is used to generate a digital human background based on user information and a determined digital human image; The generating of the digital human background based on the user information and the determined digital human image specifically includes: Determine whether the user's historical preference information is pre-stored in the system; If the user's historical preference information is pre-stored in the system, the user's identity information, the user's current scene information, the user's historical preference information and the determined digital human image are input into the first digital human background generation model, and the digital human background is generated based on the first digital human background generation model; The first digital human background generation model is: ; in, is the multidimensional vector of the output digital human background, the number of dimensions is , is the number of terms to be summed, is the number of layers, , , , , and are weights, and are all bias terms, To determine the digital human image vector The number of dimensions, Digital human image vector No. dimensions, is the user's identity information vector The number of dimensions, is the user's identity information vector No. dimensions, is the user's current scene information vector The number of dimensions, is the user's current scene information vector No. dimensions, is the user's historical preference information vector The number of dimensions, is the user's historical preference information vector No. Dimensions; If the user's historical preference information is not pre-stored in the system, the user's identity information, the user's current scene information and the determined digital human image are input into a second digital human background generation model, and a digital human background is generated based on the second digital human background generation model; The second digital human background generation model is: ; in, is the multidimensional vector of the output digital human background, the number of dimensions is , is the number of convolution operations, is the number of transformer operations, is the digital human image vector, is the user's identity information vector, is the user’s current scene information vector, is the weight, is the bias term; The second information acquisition module is used to acquire the voice interaction information input by the user and capture the user's action information in real time; A real-time interaction module, which is used to interact with the user in real time based on the voice interaction information input by the user; A digital human action determination module is used to determine the interactive action of the digital human in response to the user based on the action information of the user.

Citation Information

Patent Citations

  • Digital interaction method and system based on artificial intelligence, and medium

    CN117348736A

  • Virtual digital human driving method and device, equipment and medium

    CN117370605A

  • Digital human generation method and system, and digital human image generation device and driving method of device

    WO2024155132A1