An Active AI Digital Human Interaction Method and System Based on Visual Analysis

Through visual analysis combined with detection model methods, the accurate perception and real-time response of AI digital people to user emotions and actions is achieved, and the problems of passiveness and lag in the interaction of AI digital people in the existing technology are solved, which significantly improves the degree of user experience and interaction intelligence.

CN119806332BActive Publication Date: 2025-07-01HANGZHOU ZHILUO TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510266423.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-01
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Existing AI digital people have problems such as passivity, response lag and single-dimensional analysis in user interaction, resulting in unnatural and low intelligence.

Method used

Through the active AI digital human interaction method based on visual analysis, visual analysis combined with detection models can achieve accurate perception and real-time response to user emotions, actions and environment. The method includes batch extraction and feature extraction of image materials of preset data sources, determining training detection parameters and detection models, collecting and parsing user action and picture information in real time, and dynamically judging interaction and farewell process.

Benefits of technology

It significantly improves the intelligence, real-time and user experience of AI digital people, ensures the smoothness and integrity of the interactive process, and adapts to the interaction needs of diverse scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806332B_ABST
    Figure CN119806332B_ABST
Patent Text Reader

Abstract

The present invention provides an active AI digital human interaction method and system based on visual analysis, belonging to the field of artificial intelligence technology, including: Step 1: Extract image materials of a preset data source in batches, extract features from the image materials of each batch, determine training detection parameters based on the extracted features, and then determine a detection model; Step 2: Real-time collect the picture of the target area based on a preset collection device, analyze the picture in combination with the detection model, and determine whether to execute the interaction process based on the picture analysis result; Step 3: Execute the interaction process, obtain and analyze the information consulted by the user, and at the same time, real-time collect the user's actions based on the preset collection device and the detection model; Step 4: Analyze the user's actions, and then generate a solution plan in combination with the analysis result; Step 5: Obtain the real-time picture analysis result, and then determine whether to execute the farewell process. The present invention improves the intelligence, real-time performance and user experience of the interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an active AI digital human interaction method and system based on visual analysis. Background Art

[0002] With the rapid development of artificial intelligence technology, as an important carrier of human-computer interaction, AI digital humans have been widely used in many fields such as education, medical care, entertainment, and customer service. Through multi-modal interactions such as voice, vision, and actions, AI digital humans can provide users with immersive interaction experiences.

[0003] However, in the prior art, the interaction ability between AI digital humans and users still has obvious limitations: current AI digital humans mostly rely on active triggers from users, such as starting the interaction process through voice commands or key instructions, and cannot achieve real-time environmental perception and active response, which greatly limits the naturalness and intelligence of the user experience. When analyzing emotional characteristics such as users' expressions, actions, and tones, existing AI digital humans have problems with low accuracy. Especially in the fusion analysis of multi-modal information, there is a lack of efficient algorithm support, resulting in inaccurate understanding of user intentions, and the interaction process appears mechanical and stereotyped. In complex scenarios, AI digital humans cannot dynamically adjust interaction strategies according to user behaviors or situations. For example, in the face of users' emotional fluctuations or environmental changes, the responses of existing systems are often lagged or mismatched, thus affecting the interaction effect and user satisfaction.

[0004] Therefore, there is an urgent need for an active AI digital human interaction method based on visual analysis to achieve accurate perception and real-time response to users' emotions, actions, and environments, improve the initiative, adaptability, and interaction effect of AI digital humans, and meet the interaction needs of users in diverse scenarios. Summary of the Invention

[0005] The present invention provides an active AI digital human interaction method and system based on visual analysis, which realizes active AI digital human interaction through visual analysis in combination with a detection model, solves the problems of passive interaction mode, lagged response, and single-dimensional analysis in the prior art, and significantly improves the intelligence, real-time performance, and user experience of the interaction by accurately extracting image features, analyzing real-time images and user actions, and dynamically judging the interaction and farewell processes, ensuring the smoothness and integrity of the interaction process.

[0006] The present invention provides an active AI digital human interaction method based on visual analysis, including:

[0007] Step 1: Extract image materials of a preset data source in batches, extract features from the image materials of each batch, determine training detection parameters based on the extracted features, and further determine a detection model;

[0008] Step 2: Based on the preset acquisition equipment, the images of the target area are collected in real time, and the images are analyzed in combination with the detection model. Whether to execute the interaction process is determined based on the image analysis results;

[0009] Step 3: Execute the interaction process, obtain the information consulted by the user for analysis, and at the same time, the user's actions are collected in real time based on the preset acquisition equipment and the detection model;

[0010] Step 4: Analyze the user's actions, and then generate a solution plan in combination with the analysis results;

[0011] Step 5: Obtain the real-time image analysis results, and then determine whether to execute the farewell process.

[0012] Preferably, the image materials of the preset data source are extracted in batches, the feature extraction is performed on the image materials of each batch, and the training detection parameters are determined based on the extracted features, and then the detection model is determined, including:

[0013] The image materials of the preset data source are extracted in batches, the image materials of each batch are classified and labeled, and then several groups of image materials are determined;

[0014] Based on the first preset algorithm, several preset types of features are extracted from each group of image materials, and then several first feature sets are determined;

[0015] Based on the second preset algorithm, each first feature set is expanded, and then several second feature sets are determined;

[0016] Based on the preset learning algorithm, several third features are extracted from each second feature set, and then the third features corresponding to each second feature set are determined as the third feature set;

[0017] All the third feature sets are used as the training detection parameters, and the detection model is trained based on the training detection parameters, and then the detection model is determined.

[0018] Preferably, the batch extraction of the image materials of the preset data source includes:

[0019] Evaluate the image materials of each preset data source, and determine the stratification coefficient of each image material of each preset data source;

[0020] Based on the stratification coefficient of each image material of each preset data source and the preset coefficient-hierarchy data table, determine the data level corresponding to each image material of each preset data source, and then divide the image materials of each preset data source into several data levels;

[0021] Determine a number of sampling windows based on each data level of each preset data source and a quantum random number generator, and determine the scale of each sampling window;

[0022] Based on the scale of the sampling window of each data level of each preset data source and a preset scale-cell database, determine the size of the acquisition unit of each sampling window of each data level of each preset data source;

[0023] Perform batch extraction of image materials of each data level of each preset data source based on the size of the acquisition unit of each sampling window of each data level of each preset data source.

[0024] Preferably, evaluate the image materials of each preset data source to determine the layering coefficient of each image material of each preset data source, including:

[0025] Evaluate the image materials of each preset data source, and determine the layering coefficient of each image material of each preset data source based on the evaluation result:

[0026] ;

[0027] Wherein, is the layering coefficient of the i-th image material of the u-th preset data source, is the total number of image features of the i-th image material of the u-th preset data source, is the g-th eigenvalue of the i-th image material of the u-th preset data source, is the feature mean value of the i-th image material of the u-th preset data source, is the weight of the g-th feature of the i-th image material of the u-th preset data source, is the dynamic adjustment factor of the g-th feature of the i-th image material of the u-th preset data source, is a preset exponential amplification factor, is a preset smoothing factor, is a preset regularization factor.

[0028] Preferably, based on a preset acquisition device, the picture of the target area is collected in real time, the picture is analyzed in combination with a detection model, and it is determined whether to execute an interaction process based on the picture analysis result, including:

[0029] Lock the target area based on a dynamic area detection algorithm, and collect the picture of the target area in real time based on a preset acquisition device;

[0030] Preprocess the collected picture, and then analyze the picture collected in real time frame by frame based on the detection model to obtain the collected user information and judge whether the user actively sends an interaction signal;

[0031] Process the collected user information to determine the amount of user information;

[0032] If an interactive signal actively sent by the user is detected or the amount of user information is greater than or equal to the preset information threshold, enter the interactive process; if no interactive signal is detected or the amount of user information is less than the preset information threshold, continue to monitor the screen.

[0033] Preferably, processing the collected user information to determine the amount of user information includes:

[0034] Extract parameters from the collected user information to obtain relevant parameters of the amount of user information;

[0035] Determine the amount of user information based on the relevant parameters of the amount of user information, including:

[0036] ; where is the amount of user information, N is the number of users in the screen, is the number of time frames for action collection, is the position of the key point of the user's body in the th frame, is the position of the key point of the user's body in the th frame, is the number of user expression features in the screen, is the intensity value of the jth user expression feature, is the preset conversion coefficient of the user action amplitude, is the preset conversion coefficient of the user expression intensity.

[0037] Preferably, parsing the user action and then generating a solution plan based on the analysis result includes:

[0038] Parsing the user action and then generating a solution plan based on the analysis result includes:

[0039] Perform emotional analysis on the user action to obtain the emotional vectors of several users;

[0040] Perform interactive intention analysis on the user action and determine the interactive intention data of the user in combination with historical interactive data;

[0041] Determine the emotional vectors of the user, the interactive intention of the user, and the analysis result as the parameters of the user's needs;

[0042] Determine the demand coefficient of the user based on the parameters of the user's needs, and then generate a solution plan in combination with the preset coefficient-solution database.

[0043] The present invention provides an active AI digital human interaction system based on visual analysis, including:

[0044] Model construction module: Batch extract image materials from a preset data source, extract features from the image materials of each batch, determine training detection parameters based on the extracted features, and then determine a detection model;

[0045] Screen parsing module: Based on a preset acquisition device, it acquires the screen of the target area in real time, parses the screen in combination with the detection model, and determines whether to execute an interaction process based on the screen parsing result;

[0046] Data acquisition module: Execute the interaction process, obtain and analyze the information consulted by the user. At the same time, it acquires the user's actions in real time based on the preset acquisition device and the detection model;

[0047] Action parsing module: Parse the user's actions, and then generate a solution plan in combination with the analysis result;

[0048] Farewell execution module: Obtain the real-time screen parsing result, and then determine whether to execute the farewell process.

[0049] Compared with the prior art, the beneficial effects of the present application are as follows:

[0050] Through visual analysis, combined with the detection model, it realizes active AI digital human interaction, solves the problems of passive interaction mode, lagging response and single-dimensional parsing in the prior art. By accurately extracting image features, real-time screen and user action parsing, and dynamically judging the interaction and farewell processes, it significantly improves the intelligence, real-time performance and user experience of the interaction, and ensures the smoothness and integrity of the interaction process. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0052] Figure 1 It is a schematic flowchart of a method for active AI digital human interaction based on visual analysis provided by an embodiment of the present invention.

[0053] Figure 2 It is a schematic structural diagram of a system for active AI digital human interaction based on visual analysis provided by an embodiment of the present invention. Detailed Embodiments

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts belong to the scope of protection of the present invention.

[0055] Embodiment 1:

[0056] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, as Figure 1 shown, including:

[0057] Step 1: Extract image materials of a preset data source in batches, extract features from each batch of image materials, determine training detection parameters based on the extracted features, and then determine a detection model;

[0058] Step 2: Real-time collect the picture of the target area based on a preset collection device, parse the picture in combination with the detection model, and determine whether to execute the interaction process based on the picture parsing result;

[0059] Step 3: Execute the interaction process, obtain and analyze the information consulted by the user. At the same time, collect the user's actions in real time based on the preset collection device and the detection model;

[0060] Step 4: Parse the user's actions, and then generate a solution plan in combination with the analysis result;

[0061] Step 5: Obtain the real-time picture parsing result, and then determine whether to execute the farewell process.

[0062] In this embodiment, the preset data source includes static and dynamic image data, video data, scene data, and interference sample data.

[0063] In this embodiment, the training detection parameters are a set of core parameters for model optimization generated by extracting and analyzing the preset data source.

[0064] The beneficial effects of the above technical solutions: Through visual analysis combined with a detection model, active AI digital human interaction is realized, overcoming the problems of passive interaction, lagging response, and single parsing in the prior art. Through feature extraction and dynamic detection, the user's emotions and actions are accurately identified, and the interaction strategy is adjusted in real time, improving the naturalness and intelligence of the interaction. Combining data parsing, full-process intelligent judgment from interaction startup to farewell is realized, significantly improving the real-time performance, accuracy, and user satisfaction of the interaction, adapting to multi-scene requirements, and ensuring the high efficiency, smoothness, and personalization of the interaction process.

[0065] Embodiment 2:

[0066] An embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, which extracts image materials from a preset data source in batches, extracts features from the image materials of each batch, determines training detection parameters based on the extracted features, and then determines a detection model, including:

[0067] Extract the image materials from the preset data source in batches, classify and label the image materials of each batch, and then determine several groups of image materials;

[0068] Extract several preset types of features from each group of image materials based on a first preset algorithm, and then determine several first feature sets;

[0069] Expand each first feature set based on a second preset algorithm, and then determine several second feature sets;

[0070] Extract several third features from each second feature set based on a preset learning algorithm, and then determine the third features corresponding to each second feature set as a third feature set;

[0071] Use all the third feature sets as training detection parameters, train a detection model based on the training detection parameters, and then determine the detection model.

[0072] In this embodiment, the first preset algorithm is used to extract several preset types of features from each group of image materials, which is the initial stage of feature extraction. The focus is to capture the basic and clear features in the image. For example, the contour features in the image are extracted using an edge detection algorithm (such as Canny edge detection).

[0073] In this embodiment, the second preset algorithm is used to expand the first feature set to generate a higher-dimensional and richer feature set. For example, the local key point features are extracted using the SIFT (Scale-Invariant Feature Transform) algorithm, and at the same time, rotation and scaling transformations are applied to expand the data. Application scenario: For the same group of image materials, an extended feature set with multiple angles and different lighting conditions is generated.

[0074] In this embodiment, a preset learning algorithm is used to screen key features from the second feature set and form the final third feature set. This step is a process of feature reduction and optimization, ensuring that the features of the training model are efficient and representative. Features: Use machine learning methods (such as principal component analysis, feature selection algorithms) to screen features, extract features with high discriminative ability, and reduce the interference of irrelevant information. Example: Algorithm type: Principal component analysis (PCA): Reduce the dimension of the high-dimensional feature set and extract the most important principal components. Recursive feature elimination (RFE): Screen the features with the greatest contribution by repeatedly training the model. Application scenario: For the second feature set containing texture, color, and shape features, use PCA to extract the key features that can best distinguish user actions (such as waving, nodding) to form the third feature set.

[0075] The beneficial effects of the above technical solution: By batch extracting image materials, layer by layer extracting and expanding the feature set, and finally generating efficient training detection parameters, the performance of the detection model is optimized. Compared with the prior art, this method solves the problems of insufficient image feature extraction and low model detection accuracy, and can effectively improve the integrity of feature extraction and the detection accuracy of the model, and is applicable to active AI digital human interaction in complex scenarios.

[0076] Embodiment 3:

[0077] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, which batch extracts image materials from a preset data source, including:

[0078] Evaluate the image materials of each preset data source to determine the layering coefficient of each image material of each preset data source;

[0079] Based on the layering coefficient of each image material of each preset data source and the preset coefficient - level data table, determine the data level corresponding to each image material of each preset data source, and then divide the image materials of each preset data source into several data levels;

[0080] Based on each data level of each preset data source and the quantum random number generator, determine several sampling windows and determine the scale of each sampling window;

[0081] Based on the scale of the sampling window of each data level of each preset data source and the preset scale - unit database, determine the size of the acquisition unit of each sampling window of each data level of each preset data source;

[0082] Based on the size of the acquisition unit of each sampling window of each data level of each preset data source, batch extract the image materials of each data level of each preset data source.

[0083] In this embodiment, each image material is divided into a specific data level (such as high level, medium level, low level) according to a preset coefficient - level data table by evaluating its characteristics (such as clarity, resolution, color complexity, etc.). Materials at different levels may have different importance or analysis depths. For example, assume there is a preset data source containing multiple image materials:

[0084] Image A: High resolution, rich in details (high layering coefficient), classified as high - level data; Image B: Medium resolution, moderate in information content, classified as medium - level data; Image C: Low resolution, content blurred, classified as low - level data.

[0085] In this embodiment, the sampling window is the area in the image material for data extraction, and its scale (size) is determined by the data level. The sampling window for high - level data is generally larger, and the sampling window for low - level data is smaller to capture different levels of detailed information. For example:

[0086] For Image A at the high - level data level, the sampling window may be an area of 100×100 pixels. For Image B at the medium - level data level, the sampling window may be an area of 50×50 pixels. For Image C at the low - level data level, the sampling window may be an area of 25×25 pixels.

[0087] In this embodiment, the acquisition unit is a smaller basic data unit in the sampling window, used to refine the data extraction process. The size of the acquisition unit is usually further determined according to the scale of the sampling window and the image level. For example: For Image A at the high - level data level, the sampling window is 100×100 pixels, and the acquisition unit may be a block of 10×10 pixels; for Image B at the medium - level data level, the sampling window is 50×50 pixels, and the acquisition unit may be a block of 5×5 pixels; for Image C at the low - level data level, the sampling window is 25×25 pixels, and the acquisition unit may be a block of 2×2 pixels.

[0088] Advantages of the above - mentioned technical solution: By dividing the image data level based on the layering coefficient and dynamically determining the sizes of the sampling window and the acquisition unit in combination with the quantum random number generator, efficient batch extraction of image materials is achieved. Compared with the prior art, this method solves the problems of low efficiency and insufficient layering accuracy in the traditional image data extraction process, and significantly improves the flexibility of data processing and the accuracy of the interaction model.

[0089] Embodiment 4:

[0090] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, which evaluates the image materials of each preset data source and determines the layering coefficient of each image material of each preset data source, including:

[0091] Evaluate the image materials of each preset data source, and determine the layering coefficient of each image material of each preset data source based on the evaluation results:

[0092] ;

[0093] Among them, is the layering coefficient of the i-th image material of the u-th preset data source, is the total number of image features of the i-th image material of the u-th preset data source, is the g-th eigenvalue of the i-th image material of the u-th preset data source, is the feature mean value of the i-th image material of the u-th preset data source, is the weight of the g-th feature of the i-th image material of the u-th preset data source, is the dynamic adjustment factor of the g-th feature of the i-th image material of the u-th preset data source, is the preset exponential amplification factor, is the preset smoothing factor, is the preset regularization factor.

[0094] In this embodiment, the dynamic adjustment factor is used to adjust the influence of the eigenvalue of the image material on the layering coefficient. Its function is to dynamically enhance or weaken the importance of certain features according to specific scenarios or requirements, so as to optimize the accuracy of image layering. This factor is usually set according to real-time environmental parameters, user requirements or specific algorithms to ensure more flexible and adaptable results.

[0095] In this embodiment, evaluate the image materials of each preset data source. Among them, the evaluation process and results include: determining the total number of image features: according to the analysis requirements and technical means of the image materials, clarify the number of types of features extracted, such as extracting multiple features such as color, texture, and shape, and counting their total number; obtaining the eigenvalue: for each feature, use the corresponding feature extraction algorithm (such as extracting color feature values using a color histogram, extracting texture feature values using a texture analysis algorithm, etc.) to obtain the specific eigenvalue of the i-th image material under the u-th data source; calculating the feature mean value: sum all the feature values of the i-th image material and then divide by the total number of features to obtain the feature mean value; setting the feature weight: based on the importance of the feature (such as assigning a higher weight to the features that play a key role in image classification or layering), determine the weight of each feature g through preset artificial experience setting or machine learning algorithm training; determining the dynamic adjustment factor: according to the application scenario of the image material (such as the dynamic response to certain features in real-time processing) or additional constraint conditions, determine this factor through preset rules or real-time calculation to flexibly adjust the role of the feature.

[0096] Beneficial effects of the above technical solution: By comprehensively considering parameters such as feature weights, dynamic adjustment factors, and regularization, optimizing the calculation method of the hierarchical coefficient, and achieving precise evaluation and batch extraction of image materials. Compared with the prior art, it solves the problems of inflexible feature selection and low hierarchical accuracy, and significantly improves the data processing efficiency and the intelligent level of active AI digital human interaction.

[0097] Example 5:

[0098] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, which includes: collecting the picture of the target area in real time based on a preset acquisition device, parsing the picture in combination with a detection model, and determining whether to execute an interaction process based on the picture parsing result.

[0099] Lock the target area based on the dynamic area detection algorithm, and collect the picture of the target area in real time based on a preset acquisition device;

[0100] Preprocess the collected picture, and then parse the picture collected in real time frame by frame based on the detection model to obtain the collected user information and judge whether the user actively sends an interaction signal;

[0101] Process the collected user information, and then determine the amount of user information;

[0102] If it is detected that the user actively sends an interaction signal or the amount of user information is greater than or equal to the preset information amount threshold, enter the interaction process. If no interaction signal is detected or the amount of user information is less than the preset information amount threshold, continue to monitor the picture.

[0103] In this embodiment, the dynamic area detection algorithm is a technology for real-time analysis of picture changes, used to automatically lock the target area (such as a user or an object). It quickly determines the area that needs to be concerned by analyzing the movement, color changes, or distribution of feature points in the picture, reduces interference from irrelevant information, and improves the detection efficiency. For example, in a monitoring scenario, the dynamic area detection algorithm can detect moving objects in the picture in real time, such as a person walking in a room. The algorithm locks the position of the person as the target area and ignores the static background such as furniture and walls.

[0104] In this embodiment, the preset information amount threshold is a standard used by the system to judge the richness of information in the picture. It represents the minimum amount of information required (such as the number of features, complexity, or quality) to determine whether to continue executing the interaction process. For example, assume the threshold is 10 feature points. If only a static contour is detected in the picture and the amount of information is 5, which is lower than the threshold, the system continues to monitor. If a clear face with rich facial expressions is detected and the amount of information reaches 15, it meets the threshold requirement and enters the interaction process.

[0105] In this embodiment, the interaction process is a behavioral pattern that is initiated when the system detects that the conditions are met (such as an active signal or sufficient information). It usually includes actions, conversations, or responses to interact with the user. For example, in an AI digital human system in a shopping mall, when the user waves (active signal) or the system detects a clear face and voice (meeting the information volume threshold), the system starts the interaction process, such as greeting, recommending products, or answering the user's questions.

[0106] Beneficial effects of the above technical solution: Through the dynamic region detection algorithm and frame-by-frame parsing technology, the target region is accurately locked and user information is obtained in real time. By combining the interaction signal and the information volume threshold to determine whether to enter the interaction process, the problems of low interaction efficiency and high false trigger rate in the prior art are solved, and efficient and intelligent active AI digital human interaction is achieved.

[0107] Embodiment 6:

[0108] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, which processes the collected user information and then determines the user information volume, including:

[0109] Extract parameters from the collected user information to obtain relevant parameters of the user information volume;

[0110] Determine the user information volume based on the relevant parameters of the user information volume, including:

[0111] ; where is the user information volume, N is the number of users in the picture, is the number of time frames for action acquisition, is the position of the key point of the user's body in the th frame, is the position of the key point of the user's body in the th frame, is the number of user expression features in the picture, is the intensity value of the jth user expression feature, is the preset conversion coefficient of the user action amplitude, is the preset conversion coefficient of the user expression intensity.

[0112] In this embodiment, the key point positions of the user's body refer to the coordinate information used to describe the main joints or parts of the human body in human pose recognition. These key points usually include the head, shoulders, elbows, wrists, hips, knees, ankles, etc., which can reflect the user's movements or postures. By analyzing the changes in the positions of these key points, the amplitude of the user's movements and behavioral intentions can be judged. For example, in a dance training system, the AI ​​identifies the key point positions of the dancer through visual capture. For example, the key point coordinates of the left shoulder in the 5th frame are (200, 300), and the right shoulder is (250, 300). As time goes by, the left shoulder coordinates in the 6th frame become (220, 280), and the right shoulder becomes (270, 280). From this, the system analyzes the movement trajectory of the dancer tilting to the right.

[0113] Beneficial effects of the above technical solution: By extracting the key point positions and expression features of the user's body, and combining the movement amplitude and expression intensity conversion coefficient, the user information is accurately quantified, solving the problems of insufficient comprehensiveness in analyzing user behavior and emotion and insufficient interaction accuracy in the prior art, and improving the intelligence and real-time performance of AI digital human interaction.

[0114] Embodiment 7:

[0115] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, which analyzes the user's actions and then generates a solution based on the analysis results, including:

[0116] Analyze the user's actions, and then generate a solution based on the analysis results, including:

[0117] Perform emotional analysis on the user's actions to obtain the emotional vectors of several users;

[0118] Analyze the interaction intention of the user's actions, and combine the historical interaction data to determine the user's interaction intention data;

[0119] Determine the emotional vector of the user, the interaction intention of the user, and the analysis result as the parameters of the user's needs;

[0120] Determine the demand coefficient of the user based on the parameters of the user's needs, and then generate a solution in combination with the preset coefficient-solution database.

[0121] In this embodiment, the first analysis refers to a preliminary analysis of the user's actions to extract emotion-related features in the actions, such as body language, facial expressions, etc., for generating emotional vectors. For example, the user shows smiling and slight nodding actions through the camera, and the system identifies emotional states such as "friendly" and "happy" through the first analysis;

[0122] In this embodiment, the emotion vector is multi-dimensional data used to represent the user's current emotional state. Each dimension corresponds to different emotional characteristics (such as happiness, anger, sadness, etc.), which can quantify the user's emotions. For example, when the user shows a smile and a relaxed posture, the emotion vector may be represented as [0.8 (happy), 0.2 (calm), 0.0 (angry)].

[0123] In this embodiment, the second parsing refers to identifying the user's specific interaction intention by further analyzing the user's actions and combining historical interaction data (such as the user's previous behaviors and preferences). For example, when the user points at an object with a finger and turns their head to look at the camera, and combining the user's previous Q&A records, the second parsing determines that the user may be asking "What is this?".

[0124] In this embodiment, determining the user's demand coefficient based on the parameters of the user's needs is to comprehensively calculate the user's demand coefficient according to the emotion vector, interaction intention, and other analysis parameters, which is used to measure the urgency and clarity of the user's current needs. For example, when the system analyzes the user's hasty gestures, serious expression (the emotion vector shows anxiety of 0.9), and the voice asks "Tell me quickly what the weather is like", a high demand coefficient of 0.95 is generated, prompting the system to answer first.

[0125] The beneficial effects of the above technical solutions: Through multi-layer parsing of the user's actions, extracting the emotion vector and interaction intention, combining historical data to generate user demand parameters and demand coefficients, and accurately matching the solution, it solves the problems of inaccurate recognition of the user's emotions and intentions and non-intelligent response in the prior art, and improves the naturalness and effectiveness of the AI digital human interaction.

[0126] Embodiment 8:

[0127] The embodiment of the present invention provides an active AI digital human interaction method based on visual analysis, as Figure 2 shown, including:

[0128] Model construction module: Extract image materials from the preset data source in batches, extract features from each batch of image materials, determine training detection parameters based on the extracted features, and then determine the detection model;

[0129] Picture parsing module: Based on the preset acquisition equipment, collect the picture of the target area in real time, parse the picture in combination with the detection model, and determine whether to execute the interaction process based on the picture parsing result;

[0130] Data acquisition module: Execute the interaction process, obtain the information consulted by the user for analysis, and at the same time, collect the user's actions in real time based on the preset acquisition equipment and the detection model;

[0131] Action parsing module: Parse the user's actions, and then generate a solution in combination with the analysis results;

[0132] Farewell execution module: Obtain the real-time video analysis result, and then determine whether to execute the farewell process.

[0133] Advantages of the above technical solution: Through visual analysis and combined with a detection model, active AI digital human interaction is realized, solving the problems of passive interaction mode, lagging response, and single-dimensional analysis in the prior art. By accurately extracting image features, real-time video and user action analysis, and dynamically judging the interaction and farewell processes, the intelligence, real-time performance, and user experience of the interaction are significantly improved, ensuring the smoothness and integrity of the interaction process.

[0134] Finally, it should be noted that: The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An active AI digital human interaction method based on visual analysis, characterized in that: include: Step 1: Batch extract image materials from a preset data source, perform feature extraction on each batch of image materials, determine training detection parameters based on the extracted features, and then determine the detection model; Step 2: Based on the preset acquisition equipment, the images of the target area are collected in real time, and the images are analyzed in combination with the detection model. Based on the image analysis results, it is determined whether to execute the interactive process; Step 3: Execute the interactive process and obtain the user's consultation information for analysis. At the same time, collect user actions in real time based on the preset collection equipment and detection model; Step 4: Analyze the user's actions and generate a solution based on the analysis results; Step 5: Obtain the real-time image analysis results to determine whether to execute the farewell process The batch extraction of image materials from the preset data source includes: Evaluate the image material of each preset data source to determine the stratification coefficient of each image material of each preset data source; Determine the data level corresponding to each image material of each preset data source based on the hierarchical coefficient of each image material of each preset data source and the preset coefficient-level data table, and then divide the image material of each preset data source into a plurality of data levels; Determine a number of sampling windows based on each data level of each preset data source and the quantum random number generator, and determine the scale of each sampling window; Determine the size of the acquisition unit of each sampling window of each data level of each preset data source based on the scale of the sampling window of each data level of each preset data source and a preset scale-unit database; Batch extracting image materials of each data level of each preset data source based on the size of the acquisition unit of each sampling window of each data level of each preset data source; The image material of each preset data source is evaluated to determine the stratification coefficient of each image material of each preset data source, including: The image material of each preset data source is evaluated, and the stratification coefficient of each image material of each preset data source is determined based on the evaluation result: ; in, is the layering coefficient of the i-th image material of the u-th preset data source, is the total number of image features of the i-th image material of the u-th preset data source, is the g-th eigenvalue of the i-th image material of the u-th preset data source, is the feature mean of the i-th image material of the u-th preset data source, is the weight of the g-th feature of the i-th image material of the u-th preset data source, is the dynamic adjustment factor of the g-th feature of the i-th image material of the u-th preset data source, is the preset exponential magnification factor, is the preset smoothing factor, is the preset regularization factor.

2. The active AI digital human interaction method based on visual analysis according to claim 1, characterized in that: Batch extract the image materials of the preset data source, extract features from each batch of image materials, determine the training detection parameters based on the extracted features, and then determine the detection model, including: Batch extracting image materials from a preset data source, classifying and labeling each batch of image materials, and then determining several groups of image materials; Extracting a plurality of preset types of features from each group of image materials based on a first preset algorithm, and then determining a plurality of first feature sets; Expanding each first feature set based on a second preset algorithm, thereby determining a plurality of second feature sets; Extracting a plurality of third features from each second feature set based on a preset learning algorithm, and then determining the third features corresponding to each second feature set as a third feature set; All third feature sets are used as training detection parameters, and a detection model is trained based on the training detection parameters, thereby determining the detection model.

3. The active AI digital human interaction method based on visual analysis according to claim 1, characterized in that: Based on the preset acquisition equipment, the images of the target area are collected in real time, and the images are analyzed in combination with the detection model. Based on the image analysis results, it is determined whether to execute the interactive process, including: The target area is locked based on the dynamic area detection algorithm, and the image of the target area is captured in real time based on the preset acquisition equipment; Pre-process the captured images, and then analyze the captured images frame by frame based on the detection model to obtain the collected user information and determine whether the user actively sends an interactive signal; Processing the collected user information to determine the amount of user information; If it is detected that the user actively sends an interactive signal or the amount of user information is greater than or equal to the preset information threshold, the interactive process will be entered. If no interactive signal is detected or the amount of user information is less than the preset information threshold, the screen will continue to be monitored.

4. The active AI digital human interaction method based on visual analysis according to claim 3, characterized in that: Process the collected user information to determine the amount of user information, including: Extract parameters from the collected user information, and then obtain relevant parameters of the user information volume; The amount of user information is determined based on parameters related to the amount of user information, including: ; in, is the amount of user information, N is the number of users in the screen, is the number of time frames for action collection, is the key point position of the user's body in the kth frame, is the key point position of the user's body in the k+1th frame, is the number of user expression features in the picture, is the intensity value of the jth user's expression feature, A1 is the preset user action amplitude conversion coefficient, It is the preset user expression intensity conversion coefficient.

5. The active AI digital human interaction method based on visual analysis according to claim 1, characterized in that: Analyze user actions and generate solutions based on the analysis results, including: Perform sentiment analysis on user actions to obtain sentiment vectors of several users; Analyze the interaction intention of user actions and determine the user's interaction intention data based on historical interaction data; Determine the user's emotional vector, user's interaction intention and analysis results as parameters of user needs; The user's demand coefficient is determined based on the parameters of the user's demand, and then a solution is generated in combination with a preset coefficient-solution database.

6. An active AI digital human interaction system based on visual analysis, applied to an active AI digital human interaction method based on visual analysis as claimed in any one of claims 1 to 5, characterized in that: include: Model building module: batch extract image materials from the preset data source, extract features from each batch of image materials, determine training detection parameters based on the extracted features, and then determine the detection model; Image analysis module: collects images of the target area in real time based on preset acquisition equipment, analyzes the images in combination with the detection model, and determines whether to execute the interactive process based on the image analysis results; Data collection module: executes the interactive process and obtains user consultation information for analysis. At the same time, it collects user actions in real time based on preset collection equipment and detection models; Action analysis module: analyzes user actions and generates solutions based on the analysis results; Farewell execution module: obtains real-time image analysis results and determines whether to execute the farewell process.

Citation Information

Patent Citations

  • Emotion understanding and feedback method for full-duplex interactive question and answer digital human

    CN119311119A