Digital Human Driving Method and System for Museum Interaction Combined with Large Model Technology
By analyzing visitors' expressions and adjusting digital people's voice and resource allocation, the problem of museum digital people being unable to match visitors' emotions and insufficient resource allocation is solved, and personalized experience and efficient resource management are achieved.
Patent Information
- Application Number
- CN202510300566.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing digital people in museum interaction lack real-time emotion analysis functions and cannot adjust their behavior to match visitors' feelings, resulting in a gap in experience; resource allocation strategies are not adaptable in the face of unforeseen visitor traffic, which is prone to resource waste or shortage.
By analyzing the video stream of the guest's facial, using a convolutional neural network to identify expression changes, adjust the intonation and speed of digital human voice, combined with continuous monitoring of server load and memory usage, dynamically adjust resource allocation to optimize interactive content and server resources.
Real-time matching of digital people and visitors' emotions is achieved, the naturalness and effectiveness of interaction is improved, the stable operation of the server and the efficient utilization of resources are ensured, and overload and waste are avoided.
Smart Images

Figure CN119828896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a digital human driving method and system for museum interaction combined with large model technology. Background Art
[0002] The digital human driving method for museum interaction is a technology that uses virtual agents or digital characters to interact with visitors in a museum environment. This method uses animated characters or virtual assistants to convey information, provide guided tours, and offer educational content, thereby enhancing the visitors' sense of participation and learning experience.
[0003] However, the existing technology lacks real-time emotion analysis function, and the digital human cannot adjust its behavior to better match the actual feelings of visitors, which may lead to a break in the visitors' experience and affect the participation and satisfaction. In addition, the fixed resource allocation strategy shows insufficient adaptability in the face of unforeseen visitor traffic, and is prone to resource waste or shortage, affecting performance and visitor experience. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies in the existing technology, and to propose a digital human driving method and system for museum interaction combined with large model technology.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. The digital human driving method for museum interaction combined with large model technology includes the following steps:
[0006] Analyze the facial video stream of the visitor, extract image frames, and perform grayscale processing and normalization on each frame of the image to obtain a standardized facial image; input the standardized facial image into a convolutional neural network, analyze the image features and identify the expression changes, and generate a basic emotion state index;
[0007] Based on the basic emotion state index, adjust the tone and speed of the digital human's voice response to generate an adjusted voice feature; apply the adjusted voice feature to generate optimized interaction content;
[0008] Continuously monitor the server load situation and visitor interaction events, collect the processor usage rate and memory usage, generate system performance monitoring data; based on the system performance monitoring data, evaluate the current resource usage of the server, and calculate the number of CPU cores and memory that need to be increased or decreased to generate a resource adjustment result;
[0009] According to the resource adjustment result, adjust the resource allocation of the server, and at the same time continuously monitor the processor usage rate and memory usage, and determine whether to perform resource adjustment again according to the processor usage rate and memory usage.
[0010] Preferably, the steps for obtaining the standardized facial image are as follows:
[0011] Analyze the facial video stream of the visitor, extract consecutive image frames, and generate a set of original image frames;
[0012] Based on the set of original image frames, perform grayscale processing on each frame of the image, and calculate the brightness value after grayscale processing. The calculation formula is:
[0013] ;
[0014] where represents the red, green, and blue components of the image at position , and is the brightness value after grayscale processing at this position;
[0015] Based on the brightness value after grayscale processing, perform normalization processing on each image frame to generate a standardized facial image.
[0016] Preferably, the steps for obtaining the basic emotional state index are as follows:
[0017] Train a convolutional neural network model, input the standardized facial image into the convolutional neural network model, and the convolutional neural network model detects key features of the face through multiple layers of filters to obtain facial feature data;
[0018] Based on the facial feature data, the convolutional neural network model continues to analyze the expression changes in the image, including the fine-tuning of eyebrows and the curvature of smiles, to obtain an expression recognition result;
[0019] Based on the expression recognition result, use a classifier to map different expressions to emotional states of happiness, sadness, and surprise to obtain a basic emotional state index.
[0020] Preferably, the steps for obtaining the adjusted speech features are as follows:
[0021] Extract emotion values from the basic emotional state index, including pleasure and sadness, to obtain emotion analysis data;
[0022] Based on the emotion analysis data, calculate the intonation of the digital human's speech response. The intonation calculation formula is:
[0023] ;
[0024] where E represents the pleasure extracted from the emotion analysis data, are the coefficients for adjusting the intonation respectively, and P is the adjusted intonation;
[0025] Using the emotional analysis data, calculate the speed of the digital human's voice response, and summarize the intonation and speed to obtain the adjusted voice characteristics. The speed calculation formula is:
[0026] ;
[0027] where E also represents the pleasure degree extracted from the emotional analysis data, is the coefficient for adjusting the speed, and V is the adjusted speed.
[0028] Preferably, the steps for obtaining the system performance monitoring data are:
[0029] Continuously monitor the server, track the server's load situation and visitor interaction events, and regularly collect the CPU and memory usage data of the server to obtain real-time monitoring data;
[0030] Based on the real-time monitoring data, analyze the processor usage rate and memory usage situation, and identify performance bottlenecks or abnormal usage patterns to obtain a performance analysis result;
[0031] Based on the performance analysis result, integrate and format the data to obtain the system performance monitoring data.
[0032] Preferably, the steps for obtaining the resource adjustment result are:
[0033] Extract the real-time usage rates of CPU and memory, as well as the peak usage rate and idle rate, from the system performance monitoring data to obtain comprehensive resource usage status data;
[0034] Based on the comprehensive resource usage status data, calculate the number of CPU cores and the amount of memory that need to be adjusted. The formula is:
[0035] ;
[0036] and
[0037] ;
[0038] where, and are the current CPU and memory usage rates respectively, and are the target usage rates, is the standard deviation of the CPU usage rate, and are the number of adjusted CPU cores and the amount of memory respectively;
[0039] Based on the adjusted number of CPU cores and the amount of memory, form a resource adjustment result.
[0040] Preferably, according to the resource adjustment result, adjust the resource allocation of the server, and at the same time continuously monitor the processor utilization rate and memory usage. The steps to determine whether to perform resource adjustment again according to the processor utilization rate and memory usage are as follows:
[0041] According to the resource adjustment result, increase or decrease the number of CPU cores and memory to obtain an updated server configuration;
[0042] Based on the updated server configuration, continuously track the processor utilization rate and memory usage to obtain real-time monitoring results;
[0043] Based on the real-time monitoring results, analyze the current processor utilization rate and memory usage, evaluate whether the predetermined performance target is met, and whether resource adjustment is required again.
[0044] The present invention provides a digital human driving system, including:
[0045] An expression recognition module that captures the facial video stream of the visitor, selects key frames, grayscales and normalizes the images to generate a standardized facial image;
[0046] An emotion analysis module that inputs the standardized facial image into a convolutional neural network, analyzes the image features, recognizes the expression changes, and generates a basic emotion state index;
[0047] A voice adjustment module that adjusts the tone and speed of the digital human voice according to the basic emotion state index to generate adjusted voice features;
[0048] A content generation module that applies the adjusted voice features to create interactive content that matches the visitor's emotion to generate optimized interactive content;
[0049] A resource management module that monitors the server load and visitor interaction data, collects the processor utilization rate and memory data, generates system performance monitoring data, evaluates resource usage, and calculates the CPU and memory adjustment requirements for adjustment.
[0050] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0051] Through the analysis of the visitor's facial video stream and the processing of extracting image frames, the present invention realizes the capture of expression changes, allowing the digital human to adjust the interaction method according to the visitor's instant emotion. This ability to adjust the tone and speed in real time improves the naturalness and effectiveness of the interaction, bringing a more personalized experience to the visitor. By continuously monitoring the server load and memory usage, resources can be dynamically adjusted to ensure efficient processing of visitor requests and prevent server overload. The real-time optimization of resources not only ensures the stable operation of the server but also improves energy and cost efficiency. Brief Description of the Drawings
[0052] Figure 1 This is the step schematic diagram of the present invention. Specific implementation manner
[0053] In order to make the objectives, technical results and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] Please refer to Figure 1 , the present invention provides a technical result, a digital human driving method for museum interaction combined with large model technology, including the following steps:
[0055] Analyze the facial video stream of the visitor, extract image frames, and perform grayscale processing and normalization on each frame of the image to obtain a standardized facial image; input the standardized facial image into a convolutional neural network, analyze the image features and identify the expression changes to generate a basic emotional state index;
[0056] Based on the basic emotional state index, adjust the tone and speed of the digital human's voice response to generate an adjusted voice feature; apply the adjusted voice feature to generate optimized interaction content;
[0057] Continuously monitor the server load situation and visitor interaction events, collect the processor usage rate and memory usage situation, generate system performance monitoring data; based on the system performance monitoring data, evaluate the current resource usage situation of the server, and calculate the number of CPU cores and memory that need to be increased or decreased to generate a resource adjustment result.
[0058] According to the resource adjustment result, adjust the resource allocation of the server, and at the same time continuously monitor the processor usage rate and memory usage situation, and determine whether to perform resource adjustment again according to the processor usage rate and memory usage situation.
[0059] The steps for obtaining the standardized facial image are as follows:
[0060] Analyze the facial video stream of the visitor, extract continuous image frames, and generate an original image frame set;
[0061] Based on the original image frame set, perform grayscale processing on each frame of the image, calculate the brightness value after grayscale processing, and the calculation formula is:
[0062] ;
[0063] Among them, represents the red, green, and blue components of the image at position , is the brightness value after grayscale processing at this position;
[0064] Based on the brightness values after grayscale conversion, perform normalization processing on each image frame to generate a standardized facial image.
[0065] Specifically, continuous video data of visitors' faces is obtained by placing a CCD image sensor and a digital camera with a fixed focal length lens indoors in the museum. The analog video signal output by the camera is connected to the data acquisition module for digital conversion and caching processing. The H.264 decoder is used to extract RGB format image frames frame by frame from this data stream. The pixel data of each frame is stored in a random access memory in row and column order. The image frame rate is determined to be 30 frames per second by counting the number of decoded frames and referring to the internal recorded timestamp. The pixel value array is read frame by frame from the decoded data, and the clarity and brightness of each frame of the image are quantitatively evaluated. The clarity parameters are given by the contrast threshold 100 and the average gradient threshold 5 obtained through pre-laboratory determination. These thresholds are determined by statistically analyzing the facial images of multiple visitors under the same illumination conditions and calculating the pixel intensity contrast distribution and the local edge gradient mean. Frames with a contrast lower than 100 or an average gradient lower than 5 are directly discarded. For frames with significantly overexposed brightness, the overexposure threshold 220 obtained from the actual measured data of the on-site illuminometer is used as the judgment criterion, and frames exceeding 220 are excluded. All frames in the obtained video stream are continuously processed, and the image frames that meet the above quality standards are collected as a set of continuous original image frames and stored in memory.
[0066] The benefit of the formula is that by squaring and weighting the RGB components and using specific proportional coefficients, the calculated brightness value is closer to the human eye's perception characteristics of brightness changes. Therefore, when performing image processing based on brightness (such as normalization) subsequently, a more balanced and stable brightness distribution result can be obtained;
[0067] The steps to obtain R(i, j) are as follows: Use a high-sensitivity CCD camera to capture video frames of visitors' faces under indoor lighting conditions of approximately 300 lux in the museum exhibition hall. Read the red component of the pixel (i = 200, j = 300) from the corresponding frame. After digital processing and multiple statistics, R(i = 200, j = 300) = 185 is obtained;
[0068] The steps to obtain G(i, j) are as follows: Read the green component of the corresponding pixel in the same frame under the same conditions. After multiple measurements, the stable value is G(i = 200, j = 300) = 140;
[0069] The steps to obtain B(i, j) are the same. Extract the blue component from the same pixel position, and it is measured that B(i = 200, j = 300) = 120;
[0070] 0.299, 0.587, 0.114 are fixed values from RGB to brightness conversion parameters of ITU-R BT.601 standard;
[0071] Calculation process:
[0072] First calculate the square value: ,
[0073] , ;
[0074] Molecular computing: ,
[0075] , ;
[0076] Add the parts of the numerator: ;
[0077] Denominator calculation: ;
[0078] Final calculation ;
[0079] The result shows that the pixel is in the medium-low value area within the brightness range of 0 to 255. If the result value exceeds 128, the pixel is brighter, and below 50 it is darker. The result of 52.54 is close to the lower-middle brightness level, providing reliable brightness reference data for the subsequent image normalization step.
[0080] The brightness data of all pixels in each image frame are read from the grayscale brightness value obtained in the previous step, and these brightness values are stored in a one-dimensional array structure. The brightness values of all pixels in the array are compared in a linear traversal manner, and the maximum brightness value is recorded as Lm and the minimum brightness value is recorded as Ln. These values are obtained by statistically analyzing multiple frame data under the same lighting conditions, and are generally distributed in the range of 0 to 255. Brightness values exceeding 255 are truncated to 255, and brightness values less than 0 are directly set to 0. The obtained Lm and Ln are normalized by the calculation formula The brightness of each pixel in the frame is calculated, and the calculated normalized brightness value is directly written into the data block of the corresponding pixel position. If individual abnormal brightness values exceed the reference upper and lower limits, a limiting operation is required. The upper and lower limit values are determined by statistically analyzing the facial images of multiple visitors and calculating the average and standard deviation of the brightness of each image. For example, if the upper and lower limits are 5 and 250 respectively, the brightness value exceeding 250 is reset to 250, and that less than 5 is reset to 5. After all pixels are processed, the normalized and evenly distributed brightness array data is obtained, and the array data is stored in the memory in the form of a standardized facial image.
[0081] The steps for obtaining the basic emotional state indicators are as follows:
[0082] Train a convolutional neural network model. Input the standardized facial images into the convolutional neural network model. The convolutional neural network model detects the key features of the face through multiple layers of filters to obtain facial feature data.
[0083] Based on the facial feature data, the convolutional neural network model continues to analyze the expression changes in the image, including the fine-tuning of the eyebrows and the curvature of the smile, to obtain the expression recognition result.
[0084] Based on the expression recognition result, use a classifier to map different expressions to emotional states of happiness, sadness, and surprise to obtain the basic emotional state indicators.
[0085] Specifically, collect image samples as training inputs. First, divide these image data into a training set and a validation set according to a fixed ratio. Use the learning rate of 0.001 and a total of 30,000 iteration times that have been determined according to the distribution characteristics of multiple groups of facial image features in the experimental stage for training. Convert the images into tensor form with a fixed resolution and bit depth and load them into the video memory. Use the determined convolutional kernel size of 3×3 and 64 convolutional kernels as basic parameters. Perform forward propagation calculations on each batch of data during training, record the feature maps at each intermediate layer to identify the local texture information of the face. If the accuracy obtained on the validation set is lower than the benchmark value of 0.85 obtained from the previous statistics, then adjust the learning rate to 0.0005 in the next training cycle, and monitor the loss value during 1000 consecutive iterations. If the loss value continues to be higher than the threshold of 0.2 selected in multiple preliminary experiments, then reduce the batch size of each iteration from 64 to 32 to increase the parameter update strength. The pooling operation window size and stride have been determined through diverse tests. After multiple rounds of training, record the output of the final convolutional layer and flatten it. Subsequently, connect the fully connected layer to calculate the numerical feature vector available for subsequent analysis, summarize all the feature vectors to form a feature data set, and obtain the facial feature data.
[0086] Based on the facial feature data obtained from the previous paragraph, specific dimensions are first selected from these feature data to characterize the eyebrow area and the smile arc area. These dimensions were fixed during the previous statistics on multiple groups of annotated images. The selected feature dimensions are recorded over time series and the differences between adjacent frames are calculated. If the change amplitude of a certain feature value exceeds 0.05 in multiple consecutive frames, and this 0.05 is determined by the average expression change amplitude of more than a hundred subjects through statistics, then this feature is regarded as a sensitive feature and further differential processing is performed on it to obtain the change rate. Using the average feature value in the expressionless state as the reference benchmark, the change rate is compared with the difference of this benchmark. If the difference is greater than 0.03, and this 0.03 value is determined after analyzing the expression change ranges of multiple subjects in the laboratory environment multiple times, then it is recorded that the expression feature has changed significantly in the current frame. When it is determined that the change comes from the eyebrow feature dimension, it is marked as eyebrow fine-tuning, and when it comes from the smile arc dimension, it is marked as smile feature change. Finally, the time change features of all the marked expression features are statistically analyzed and matched with the pre-constructed correspondence table between feature changes and expression categories. When the match is successful, the corresponding expression category is recorded to obtain the expression recognition result.
[0087] Based on the obtained expression recognition result, the recognition result is matched according to the pre-constructed mapping list. This mapping list is established by collecting the expression change data of multiple subjects in different situations. The corresponding emotion labels are searched in the list for various expressions in the recognition result. If it corresponds to a happy expression, it is marked as happy, if it corresponds to a sad expression, it is marked as sad, and if it corresponds to a surprised expression, it is marked as surprised. The proportions of the appearance of each emotion label in all the marked results are statistically analyzed. When the appearance proportion of a certain emotion label exceeds 0.5, and this 0.5 value is determined by selecting the median proportion from the statistics of multiple groups of emotion distributions, then this emotion is regarded as the dominant emotion and further quantified. The emotion state is converted into digital indicators by statistically analyzing the appearance frequency of each expression and the intensity of its feature value change. Finally, these digital indicators are output to obtain the basic emotion state indicators.
[0088] The steps for obtaining the adjusted voice features are as follows:
[0089] Extract the emotion values from the basic emotion state indicators, including the pleasantness and sadness, to obtain the emotion analysis data;
[0090] Based on the emotion analysis data, calculate the intonation of the digital human's voice response. The intonation calculation formula is:
[0091] ;
[0092] where E represents the pleasantness extracted from the emotion analysis data, are the coefficients for adjusting the intonation respectively, and P is the adjusted intonation;
[0093] Using the sentiment analysis data, calculate the speed of the digital human's voice response, and summarize the tone and speed to obtain the adjusted voice features. The speed calculation formula is:
[0094] ;
[0095] Here, E also represents the pleasure extracted from the sentiment analysis data. is the coefficient for adjusting the speed, and V is the adjusted speed.
[0096] Specifically, the emotion value corresponding to the expression is read from the basic emotion state index obtained in the previous step, and the emotion value is recorded in the form of an array containing a set of pleasure and sadness. First, the pleasure and sadness stored in the array are checked item by item, and the reference range of pleasure of 0.0 to 1.0 and the reference range of sadness of 0.0 to 1.0 obtained by continuous observation of the emotional states of multiple volunteers in the early stage are used as benchmarks. Each emotion value is compared with these reference ranges. If the pleasure value is higher than 0.8, the point is recorded and marked as a high pleasure interval in the subsequent steps. The 0.8 threshold is determined by collecting facial expression data of no less than 30 resident visitors when viewing the exhibition and calculating the correlation between their expression parameters and the scores of the self-assessment questionnaire. If the sadness value is between 0.4 and 0.7, it is marked as medium in the subsequent steps. The sadness interval is equal. The values of 0.4 and 0.7 are calculated by extracting the features of the volunteers who continuously show mild depression in the same batch. The high happiness and medium sadness data points marked in all the records are summarized, and then each data point is arranged in chronological order. These data points are read in sequence using a clear sequence index method, and they are matched with the expression change curve in the corresponding time period. The linear difference calculation of adjacent data points is performed at a fixed time interval of 5 to obtain the happiness change rate and sadness change rate. The time interval required for the calculation of these change rates is 5, which is selected after statistics on the fluctuation frequency of happiness changes in multiple samples to ensure that the change rate can represent the emotional fluctuation amount in a shorter period of time. All the change rate data are recorded to form a set of emotion analysis data for subsequent analysis. The emotion analysis data is the output of the final step.
[0097] The benefit of the formula is that it uses multiple parameters to participate in the operation, linking the quantified value of pleasure with each adjustment coefficient, thereby achieving dynamic adjustment of the emotional state in the process of speech feature generation;
[0098] The steps to obtain the E parameter are as follows: Based on the emotion analysis data obtained above, continuous facial expression collection and emotion self-assessment questionnaire statistics are conducted on 200 visitors in a museum environment, and the characteristics such as skin light reflectivity and facial muscle stretch are quantified into a pleasure range of 0 to 1, from which a specific sample is selected with a pleasure level of ;
[0099] The steps for obtaining the parameters are as follows: Among the previously extracted voice data of multiple visitors, measure the acoustic features of voice samples with different pleasure degree markings, and fit the relationship between the voice frequency change and the pleasure degree through linear regression to obtain a set of tone adjustment coefficients suitable for the current scenario, where 、 、 、 , these values are obtained by statistically analyzing the sound wave frequency and the corresponding pleasure degree after collecting more than a hundred pleasant voice samples in the laboratory and then performing regression to solve;
[0100] Calculation process:
[0101] Calculate : ;
[0102] Calculate : ;
[0103] Calculate : ;
[0104] Square root calculation ;
[0105] Calculate : ;
[0106] Calculate : ;
[0107] Calculate ;
[0108] Final result ;
[0109] This result shows that when the pleasure degree is 0.65, the calculated adjusted tone is 1.1652. When the result exceeds 1.0, it means that the voice tone is more inclined to the rising state compared to the neutral tone, and the higher the value, the more high-pitched the voice tone becomes. When the result is lower than 1.0, it means that the tone is more inclined to be flat or descending.
[0110] The benefit of the formula is that through multiple operations on the pleasure degree and the adjustment speed coefficient, the voice rate can be precisely adjusted in the emotional context;
[0111] The steps for obtaining the E parameter are the same as above, and the same pleasure degree of 0.65 is used as in the previous formula;
[0112] The steps for obtaining the parameters are as follows: Measure the speech rates of multiple visitors under the same experimental conditions, establish a mapping relationship between the speech rates of each speech sample and the corresponding pleasure levels, and determine the coefficient values through multiple linear regression analyses, where and and ;
[0113] Calculation process:
[0114] Calculate : ;
[0115] Calculate ;
[0116] Calculate : ;
[0117] Calculate : ;
[0118] Calculate ;
[0119] Final result ;
[0120] This result indicates that when the pleasure level is 0.65, the calculated adjusted speech rate is 1.5202. A value greater than 1.0 means the speech rate is faster than the baseline rate, and the larger the value, the higher the speaking speed. A value lower than 1.0 means the speech rate is relatively slow.
[0121] The steps for obtaining the system performance monitoring data are as follows:
[0122] Continuously monitor the server, track the server load situation and visitor interaction events, regularly collect the CPU and memory usage data of the server to obtain real-time monitoring data;
[0123] Based on the real-time monitoring data, analyze the processor usage rate and memory usage situation, identify performance bottlenecks or abnormal usage patterns to obtain performance analysis results;
[0124] Based on the performance analysis results, integrate and format the data to obtain the system performance monitoring data.
[0125] Specifically, during the continuous monitoring of the server, the deployed monitoring program is selected to periodically read the CPU activity frequency and memory occupancy of the target server. The status information interface provided inside the server is used as the data source. The CPU occupancy ratio is obtained by calling the kernel-level status query instruction at a fixed interval in the server operating system. The memory usage rate is deduced from the amount of allocated memory pages recorded in the memory mapping table storing the process running parameters. The query interval for the CPU usage rate and memory occupancy is set to 10 seconds. This 10-second interval value is determined by continuously observing at least 72 hours on multiple servers with the same configuration in high-load and idle states, collecting CPU and memory data points, calculating the average interval response speed, and selecting the time length that can ensure the synchronization between the processor load data and the changes in visitor interaction events. The CPU usage rate and memory occupancy values obtained from each query are stored in the memory space in the form of an ordered list, and each piece of data in the list is marked with a timestamp according to the time sequence. The data points with a CPU usage rate exceeding 85% are marked. This 85% is determined by statistically analyzing the historical records provided by multiple system maintenance personnel and selecting the average value of the critical CPU usage rate at which the server performance significantly deteriorated during past operations. If the memory occupancy remains above 70% in multiple consecutive queries, it is also marked. This 70% is established by statistically analyzing the memory consumption curve in diverse usage scenarios of the application system and selecting the average value at which obvious signs of slow response occur. All the marked data and unmarked data are placed in the same data structure, and the time series change trends of these data points are read and compared one by one through the index pointer. After the recording is completed, the real-time data containing CPU and memory usage information can be sorted out to obtain the real-time monitoring data.
[0126] Based on the obtained real-time monitoring data, the CPU usage rate and memory occupancy values in each group of data points are compared. First, during the execution process, the CPU usage rate is sequentially extracted and arranged in time sequence to form a sequence, and the memory occupancy values are arranged in the same way to form another sequence. The corresponding values within the same time period are selected from the two sequences for corresponding comparison. If there are three or more consecutive values in the CPU usage rate sequence higher than the previously determined 85% marked value, the corresponding memory occupancy values at these moments are extracted. If the memory occupancy values also exceed 70% at these moments, the time period is marked as a potential performance bottleneck point in the analysis. This operation is completed through simple conditional judgments. If abnormal fluctuations below the predetermined range are observed in the CPU or memory data sequences, the predetermined range is obtained by analyzing the variance and standard deviation of the historical data and selecting the values that fluctuate 5% above and below the median, then the data point is recorded and classified as an abnormal usage pattern. All the marked potential performance bottleneck points and abnormal usage pattern points are listed together, and finally, these marked values are output to clarify the performance characteristics of these time periods, obtaining the performance analysis result.
[0127] Based on the obtained performance analysis results, retrieve the CPU and memory data values involved in the above-marked performance bottleneck points and abnormal usage pattern points, sort these data values in ascending order of time, assign sequence numbers to these data points in turn using integer sequences, read the corresponding CPU and memory values for each data point respectively, output the CPU values in percentage form and the memory values in MB unit during the processing, mark the points with CPU usage rate exceeding 85% as "high load state", mark the points with memory occupancy exceeding 70% as "high memory state", add marks to the records if the data points satisfy both high load and high memory states, and gather all these sorted data into a set of normalized structures to obtain the system performance monitoring data.
[0128] The steps to obtain the resource adjustment results are as follows:
[0129] Extract the real-time usage rates, peak usage rates, and idle rates of CPU and memory from the system performance monitoring data to obtain the comprehensive resource usage status data;
[0130] Based on the comprehensive resource usage status data, calculate the number of CPU cores and the amount of memory to be adjusted. The formulas are:
[0131] ;
[0132] and
[0133] ;
[0134] Among them, and are the current CPU and memory usage rates respectively, and are the target usage rates, is the standard deviation of the CPU usage rate, and are the number of CPU cores and the amount of memory to be adjusted respectively;
[0135] Based on the adjusted number of CPU cores and the amount of memory, form the resource adjustment results.
[0136] Specifically, select the information corresponding to the CPU and memory status from the system performance monitoring data obtained in the previous steps as the input. Arrange this information in a time series format, read the CPU usage rate at each time point, and record this value in floating-point format. Similarly, read and record the memory usage rate. By performing multiple periodic queries within a continuous period of time, obtain all the CPU and memory usage data points during this period. Attach a timestamp to each data point to maintain the coherence of the time order. Use clear numerical tags to identify each record, and map these identifiers to the CPU usage rate and memory usage rate at the corresponding time points. Then, select a time window from these data, compare each data point within the window in sequence, and identify the peak and idle rates of the CPU and memory by calculating the maximum and minimum values. The peak usage rate is selected by finding the highest CPU and memory values among all the data points within this time window, and the idle rate is obtained by finding the data point with the lowest usage rate. To ensure the credibility of the peak and idle values, it is necessary to determine the length and interval of the time window. Here, the length of the time window is set to 1 hour, which is obtained by referring to the CPU and memory change characteristics of multiple sets of system logs and summarizing the experience of operation and maintenance personnel in dozens of monitoring sessions. The interval time is set to 5 minutes, which is determined by the distribution of the CPU and memory change rates recorded by multiple servers with the same configuration under different load scenarios. Arrange and associate the calculated peak and idle values with the real-time usage rates in the corresponding time series into the same structure, and number each usage rate value using a set of integer indexes to form an ordered data sequence. Extract the real-time usage rate, peak usage rate, and idle rate contained in this sequence and integrate them into comprehensive resource usage status data that can be directly read.
[0137] The advantage of the formula is that it can perform quantitative adjustment on the number of CPU cores and the amount of memory by using parameters such as the current CPU and memory usage rates, the target usage rate, and the standard deviation for calculation.
[0138] The steps to obtain the parameter are as follows: Read the current CPU usage rate from the comprehensive resource usage status data obtained previously. The measured CPU usage rate at this time is 0.78, and this value is obtained based on the continuous collection and statistics of the CPU usage rate in the on-site monitoring data.
[0139] The steps to obtain the parameter are as follows: The operation and maintenance personnel determine the target CPU usage rate to be 0.6 according to the actual operation requirements and the statistics of historical operation data. This value is determined by analyzing the average CPU usage values of the system under high concurrency and normal load in dozens of times.
[0140] The steps to obtain the parameter are as follows: Calculate the standard deviation of the CPU usage rate in the same historical monitoring dataset. After continuous recording and statistics for one week, the standard deviation of the CPU usage rate is approximately 0.1, which is obtained by taking the square root after calculating the standard variance of the collected CPU usage rate sample data;
[0141] The steps to obtain the parameter are as follows: Extract the current memory usage rate from the previously obtained comprehensive resource usage data. After recording and analysis, the memory usage rate at this time is 0.72, which is the memory occupancy ratio at the current moment selected after multiple rounds of actual measurements;
[0142] The steps to obtain the parameter are as follows: After multi-day statistics and control experiments on the memory usage of the operating environment, the operation and maintenance personnel determine that the target memory usage rate is 0.5, which is set by selecting the average value of the memory usage during the period with better performance;
[0143] Calculation process:
[0144] For Formula:
[0145] Calculate the numerator : 0.78 - 0.6 = 0.18;
[0146] Calculate the denominator : It is 0.1;
[0147] ;
[0148] ;
[0149] ;
[0150] For Formula:
[0151] Calculate : 0.72 - 0.5 = 0.22;
[0152] ;
[0153] ;
[0154] ;
[0155] This result indicates that when the current CPU usage rate is 0.78, the target is 0.6, and the standard deviation is 0.1, the calculated , a value greater than 1 indicates that the number of CPU cores needs to be increased by approximately 1.62 cores, which is usually rounded up to 2 cores according to the actual hardware conditions; when the current memory usage rate is 0.72 and the target is 0.5, it is calculated that , a smaller value indicates that the amount of memory needs to be increased by approximately 0.0344 units, and the memory amount can be increased with the minimum step according to the actual memory increment unit.
[0156] Based on the calculated adjusted number of CPU cores and memory amount, select the target values for actual execution of resource adjustment from the parameters summarized in the previous step, and use the and values are processed through integerization rules. For it is rounded up to 2 according to the minimum adjustable unit of 1 CPU core that the hardware can operate. For it is rounded up according to the minimum memory module capacity increment of 256MB and added to 256MB. All values involved in the execution process are obtained by associating the statistically recorded hardware device parameters. These hardware parameters are recorded after actual measurement of CPUs and memories of different brands and models in the computer room. The processed number of CPU cores and memory increment are summarized and incorporated into an integrated dataset, and the final adjustment parameters of the CPU and memory are marked with a clear index number. Compare it with the previously calculated target usage rate to check whether the difference between the two is lower than the pre-selected stability threshold. This threshold of 0.05 is determined by the average difference in the overall system performance after multiple adjustments. When it is confirmed that the difference is lower than 0.05, the final resource adjustment data can be recorded to obtain the resource adjustment result.
[0157] According to the resource adjustment result, adjust the resource allocation of the server, and at the same time continuously monitor the processor usage rate and memory usage. The steps to determine whether to perform resource adjustment again based on the processor usage rate and memory usage are as follows:
[0158] According to the resource adjustment result, increase or decrease the number of CPU cores and memory to obtain an updated server configuration;
[0159] Based on the updated server configuration, continuously track the processor usage rate and memory usage to obtain real-time monitoring results;
[0160] Based on the real-time monitoring results, analyze the current processor usage rate and memory usage, and evaluate whether the predetermined performance target is met and whether resource adjustment needs to be performed again.
[0161] Specifically, according to the resource adjustment results obtained previously, read the parameters of the number of CPU cores and memory capacity that need to be increased or decreased from them. Use these parameters as the resource configuration information to be executed, and modify the target virtual machine by accessing the configuration interface provided in the virtualization management platform. First, log in to the virtualization management console and retrieve the unique identifier of the target virtual machine. Find the corresponding configuration option in the virtualization management interface through this identifier. Input the integer increase or decrease value obtained previously for the CPU core allocation parameter. For example, if it is necessary to increase from the existing 2 CPU cores to 4 CPU cores, then modify the CPU core value to 4 in the setting interface and apply the configuration. Similarly, adjust the memory capacity with reference to the memory increase or decrease value obtained previously. If it is necessary to increase 256MB of memory, then increase the allocation value of 256MB in the memory parameter column and submit the change. After submission, the virtualization management platform will dynamically allocate the underlying virtualization layer resources in the background, update the CPU and memory parameters of the virtual machine to the new values. There is no need to take the virtual machine offline and restart after the update operation. After completing the configuration modification, record the updated number of CPU cores and memory capacity and perform a comparison check. If the set value has taken effect successfully, record the completion of this adjustment operation. Finally, obtain the updated server configuration.
[0162] Based on the updated server configuration, regularly query the processor usage rate and memory usage data of the virtual machine that already has the newly added CPU cores and memory capacity. Send the query instruction to the virtual machine through the status interface provided by the virtualization management platform. Repeat the query instruction within a time period with an interval of 5 seconds. Each query can obtain the current CPU usage rate and memory usage value. Record the obtained CPU and memory occupancy rates in the form of floating-point numbers, and attach a timestamp to each record. Arrange these data points in an orderly manner to form a time series. When the CPU usage rate exceeds the 80% reference value statistically obtained from historical data multiple times, mark the corresponding time point as the CPU high-load interval. This 80% reference value is determined by selecting the most frequently occurring high-value segment from the CPU occupancy situations in multiple batches of running samples. When the memory usage rate exceeds the preset 70% threshold multiple times, mark the corresponding time point as the memory high-usage interval in the record. This 70% threshold is obtained from the experience statistics of multiple operation and maintenance personnel in different load scenarios. Associate the marking information with the original data points for storage. Traverse the entire time series to count the occurrence times of the high-load and high-usage intervals and record the results. Finally, generate a real-time monitoring result containing the latest usage rates and marking information.
[0163] Based on the real-time monitoring results, the CPU and memory usage rates in the current period are selected from the recorded data sequence for comparison, and the CPU and memory values at each time point are compared with the previously set target usage ranges. These target ranges are obtained by statistics of the operating data of the past two weeks. For example, the target CPU usage range is between 0.6 and 0.7, and the target memory usage range is between 0.4 and 0.5. If the CPU usage rate is found to be continuously higher than 0.7 in multiple consecutive queries, the information is recorded. If the CPU usage rate is high in a certain number of repeated queries, the situation is identified as a potential readjustment signal. Similarly, if the memory usage rate is higher than 0.5 in multiple queries, it is marked as excessive consumption of memory resources. The number of deviations is calculated and compared with the critical number determined in the statistics. If the critical number is exceeded, it indicates that resource allocation may need to be adjusted again. If all recorded points are within the target range, it means that the current resource allocation state can be maintained without reallocating resources. According to the final statistical results, it is determined whether the current system meets the predetermined performance goals and whether further resource adjustments are needed in the future.
[0164] The present invention provides a digital human driving system, comprising:
[0165] The expression recognition module captures the visitor's facial video stream, selects key frames, grayscales and normalizes the image, and generates a standardized facial image;
[0166] The emotion analysis module inputs the standardized facial images into the convolutional neural network, analyzes the image features, identifies expression changes, and generates basic emotional state indicators;
[0167] The voice adjustment module adjusts the tone and speed of the digital human voice according to the basic emotional state indicators and generates the adjusted voice features;
[0168] The content generation module applies the adjusted voice features to create interactive content that matches the visitor's emotions and generates optimized interactive content;
[0169] Resource management module, monitors server load and visitor interaction data, collects processor usage and memory data, generates system performance monitoring data, evaluates resource usage, and calculates CPU and memory adjustment requirements for adjustment.
[0170] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical results of the present invention still falls within the protection scope of the technical results of the present invention.
Claims
1. A digital human-driven method for museum interaction combined with large model technology, characterized in that, Include the following steps: Analyze the facial video stream of the visitor, extract image frames, and perform grayscale processing and normalization on each frame of the image to obtain a standardized facial image; Input the standardized facial image into a convolutional neural network, analyze the image features and identify expression changes to generate a basic emotional state index; Based on the basic emotional state index, adjust the tone and speed of the digital human's voice response to generate an adjusted voice feature; Apply the adjusted voice feature to generate optimized interaction content; Continuously monitor the server load situation and visitor interaction events, collect the processor usage rate and memory usage, and generate system performance monitoring data; Based on the system performance monitoring data, evaluate the current resource usage of the server, and calculate the number of CPU cores and memory that need to be increased or decreased to generate a resource adjustment result; According to the resource adjustment result, adjust the resource allocation of the server, and at the same time continuously monitor the processor usage rate and memory usage, and determine whether to perform resource adjustment again based on the processor usage rate and memory usage; The steps for obtaining the adjusted voice feature are: Extract the emotion values from the basic emotional state index, including the degree of pleasure and the degree of sadness, to obtain emotion analysis data; Based on the emotion analysis data, calculate the tone of the digital human's voice response. The tone calculation formula is: ; Among them, represents the pleasure degree extracted from the sentiment analysis data, are respectively the coefficients for adjusting the intonation, is the adjusted intonation; Use the emotion analysis data to calculate the speed of the digital human's voice response, and summarize the tone and speed to obtain the adjusted voice feature. The speed calculation formula is: ; Among them, also represents the pleasure degree extracted from the sentiment analysis data, is the coefficient for adjusting the speed, is the adjusted speed.
2. The digital human driving method for museum interaction combining large model technology according to claim 1, wherein The steps for obtaining the standardized facial image are: Analyze the facial video stream of the visitor, extract consecutive image frames, and generate a set of original image frames; Based on the set of original image frames, perform grayscale processing on each frame of the image, and calculate the brightness value after grayscale processing. The calculation formula is: ; Among them, represents the red, green, and blue components of the image at position , and is the brightness value after grayscale conversion at this position; Based on the brightness value after grayscale processing, perform normalization processing on each image frame to generate a standardized facial image.
3. The digital human driving method for museum interaction combining large model technology according to claim 1, characterized in that, The steps for obtaining the basic emotional state index are: Train a convolutional neural network model, input the standardized facial image into the convolutional neural network model, and the convolutional neural network model detects the key features of the face through multiple layers of filters to obtain facial feature data; Based on the facial feature data, the convolutional neural network model continues to analyze the expression changes in the image, including the subtle adjustment of the eyebrows and the curvature of the smile, to obtain an expression recognition result; Based on the expression recognition result, use a classifier to map different expressions to emotional states of happiness, sadness, and surprise to obtain a basic emotional state index.
4. The digital human driving method for museum interaction combined with large model technology according to claim 1, wherein The steps for obtaining the system performance monitoring data are: Continuously monitor the server, track the server load situation and visitor interaction events, and regularly collect the CPU and memory usage data of the server to obtain real-time monitoring data; Based on the real-time monitoring data, analyze the processor usage rate and memory usage, and identify performance bottlenecks or abnormal usage patterns to obtain a performance analysis result; Based on the performance analysis result, integrate and format the data to obtain system performance monitoring data.
5. The digital human driving method for museum interaction combined with large model technology according to claim 1, wherein The steps for obtaining the resource adjustment result are: Extract the real-time utilization rates, peak utilization rates, and idle rates of the CPU and memory from the system performance monitoring data to obtain comprehensive resource usage status data; Based on the comprehensive resource usage status data, calculate the number of CPU cores and the amount of memory that need to be adjusted. The formulas are: ; and ; Wherein, and are the current CPU and memory usage rates respectively, and are the target usage rates, is the standard deviation of the CPU usage rate, and are the adjusted number of CPU cores and the amount of memory respectively; Based on the adjusted number of CPU cores and the amount of memory, form a resource adjustment result.
6. The digital human driving method for museum interaction combining large model technology according to claim 1, characterized in that According to the resource adjustment result, adjust the resource allocation of the server. At the same time, continuously monitor the processor utilization rate and memory usage, and the steps to determine whether to perform resource adjustment again based on the processor utilization rate and memory usage are: According to the resource adjustment result, implement the increase or decrease of the number of CPU cores and memory to obtain an updated server configuration; Based on the updated server configuration, continuously track the processor utilization rate and memory usage to obtain real-time monitoring results; Based on the real-time monitoring results, analyze the current processor utilization rate and memory usage, and evaluate whether the predetermined performance target is met and whether resource adjustment is needed again.
7. The digital human driving system for museum interaction using the large model technology according to any one of claims 1-6, characterized in that, Include: An expression recognition module that captures the facial video stream of the visitor, selects key frames, grayscales and normalizes the images to generate a standardized facial image; An emotion analysis module that inputs the standardized facial image into a convolutional neural network, analyzes the image features, recognizes the expression changes, and generates basic emotion state indicators; A voice adjustment module that adjusts the tone and speed of the digital human voice according to the basic emotion state indicators to generate adjusted voice features; A content generation module that applies the adjusted voice features to create interactive content that matches the visitor's emotion and generates optimized interactive content; A resource management module that monitors the server load and visitor interaction data, collects the processor utilization rate and memory data, generates system performance monitoring data, evaluates resource usage, and calculates the CPU and memory adjustment requirements for adjustment.
Citation Information
Patent Citations
AI digital human interaction method and system based on emotion recognition
CN118519538A
Robot facial expression regulation and control system and method based on internal projection technology
CN118700178A
Server resource utilization rate improving method, system, device, medium and product
CN119292776A
Digital human posture adjusting method and device combined with image recognition
CN119473209A