Power marketing customer behavior identification method and system based on multi-modal analysis
By using multimodal data collection and feature fusion, the problem of accuracy in customer behavior recognition in complex scenarios has been solved, achieving efficient optimization of business hall resources and improvement of customer service.
Patent Information
- Application Number
- CN202510882086.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Existing customer behavior recognition systems perform poorly in complex scenarios. Single-modal data is greatly affected by noise, lighting, and individual customer differences. The asynchronous temporal dimension of multimodal data leads to inaccurate feature alignment. They cannot adapt to different environments and customer groups and lack long-term behavioral trend modeling, resulting in unstable prediction results.
By deploying multimodal data acquisition equipment, time synchronization alignment and cross-modal feature extraction are performed. Combined with the cross-modal Transformer feature alignment model, feature fusion is carried out using voice, video and environmental data. Deep reinforcement learning is used to optimize model parameters, predict subsequent customer behavior, and optimize business processes.
It achieves high-precision customer behavior recognition in complex environments, improves the efficiency of service halls and customer experience, automatically adjusts resource allocation, reduces waiting time, increases the rate of self-service transactions, and enhances the stability of voice recognition and the robustness of behavior recognition.
Smart Images

Figure CN120995064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electricity marketing technology, specifically to a method and system for identifying electricity marketing customer behavior based on multimodal analysis. Background Technology
[0002] Most existing customer behavior recognition systems rely solely on single-modal data, such as voice emotion analysis or video object detection. While this approach works in simple scenarios, its effectiveness significantly decreases in complex environments (such as customer service settings in a sales office) because single-modal data can be affected by factors like noise, lighting conditions, and individual customer differences. For example: When relying solely on speech recognition, background noise can distort speech features and reduce the accuracy of emotion recognition.
[0003] When relying solely on video behavior analysis, it is difficult to accurately infer a customer's true intentions from their body language, especially when the customer is waiting or anxious, as their posture may be unstable, making it difficult to accurately determine the customer's business needs.
[0004] In existing multimodal behavior analysis methods, the time dimensions of different data sources (such as voice and video) are not synchronized, leading to inaccurate feature alignment. For example: Speech data sampling rates are typically high (e.g., 16kHz), while video frame rates are generally 30fps. This difference can lead to data misalignment issues during model fusion.
[0005] Existing technologies typically employ linear interpolation or fixed window alignment for time synchronization, but these methods cannot accurately align complex behavioral features, resulting in significant errors in behavioral analysis.
[0006] Most current customer behavior prediction methods use single time series models (such as LSTM or simple rule matching), but this method has a significant problem with insufficient generalization ability: The existing model cannot adapt to different business hall environments. Different business halls have different customer groups and business needs, and the existing model often requires manual adjustment of parameters, which cannot dynamically adapt to environmental changes.
[0007] The inability to accurately predict customers' next actions, relying solely on short-term behavioral patterns and lacking long-term behavioral trend modeling, leads to unstable prediction results.
[0008] The environment of a sales office (such as temperature, humidity, noise, and air quality) has a significant impact on customer behavior. For example, a high-noise environment can cause changes in customers' speech rate and reduced speech clarity, but current technologies often fail to dynamically adjust the recognition model, leading to a decline in service quality. Furthermore, existing methods lack adaptive adjustments to changes in customer emotions and fail to incorporate multimodal data to optimize service strategies. Summary of the Invention
[0009] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for identifying customer behavior in electricity marketing based on multimodal analysis, comprising: Deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data on the customer's voice, video, and environment; The collected multimodal data is time-synchronized and aligned, and features are extracted using a cross-modal feature alignment method. The fused customer behavior vectors are input into the customer behavior recognition model for real-time analysis to determine customer business needs, predict subsequent customer behavior, and optimize business processes based on customer behavior data.
[0010] As a preferred embodiment of the electricity marketing customer behavior recognition method based on multimodal analysis described in this invention, the method includes: arranging a microphone array to acquire a mixed audio signal containing the target customer's voice and environmental noise; performing beamforming processing on the acquired audio signal; calculating the sound source direction based on the time difference, phase difference, and amplitude changes of the signals received by different microphones; and adjusting the pickup weights of the microphone array. Ultrasonic sensors are deployed to emit ultrasonic signals and receive echo signals reflected from environmental objects and the human body. Based on the time of flight of the ultrasonic signals, the time difference of echoes at different locations is calculated, and the direction of the sound source identified by beamforming technology is used to determine the location of the customer's voice. Frequency modulation analysis is performed on the ultrasonic signals to distinguish between the customer's voice echo and the environmental noise echo, and the positioning data is matched with the pickup direction of the microphone array. Short-time Fourier transform (SFT) is performed on the target customer's speech signal after beamforming and voice location detection to decompose the time-frequency characteristics of the audio signal. The spectrum after the SFT is mapped using a Mel filter bank, and Mel spectral features are extracted. Based on the extracted Mel spectrum, the spectral envelope, frequency band energy distribution, and formant features of the speech signal are calculated to obtain the timbre characteristics of the speech signal. The time-domain characteristics of the speech signal are analyzed by calculating the speech rate, energy fluctuations, pause intervals, and fundamental frequency trajectory to form the speech temporal characteristics. Based on the extracted Mel spectrum features and speech temporal features, an emotion classification model is constructed. The spectral variation trend, speech rate pattern and energy distribution of the speech signal in different time windows are analyzed, the dynamic changes of pitch, intensity and rhythm are calculated, and feature parameters that can characterize emotional state are extracted. The classification model is used to analyze the feature parameters, and combined with the variation trend of temporal features, the customer's emotion category is identified, and the speech emotion recognition result is output.
[0011] As a preferred embodiment of the electricity marketing customer behavior recognition method based on multimodal analysis described in this invention, the method involves: deploying camera equipment to acquire a continuous video stream containing customer head movements, gestures, and body movements; performing frame segmentation on the video data; and extracting frame sequences for behavior analysis. Target detection is performed on the video frame sequence, and the positions of the customer's head, hands and torso are extracted using a human pose detection algorithm. The coordinates of the detected key points are smoothed based on the Kalman filter algorithm, and the motion trajectory between consecutive frames is analyzed by combining optical flow method to establish the motion time series of the customer's head, gestures and limb movements. The system uses cameras to acquire the three-dimensional spatial coordinates of various joints on the customer's body, performs three-dimensional modeling of the joints, and uses Euclidean distance to calculate the spatial relationship between different joints to generate a skeleton structure model. Based on time series analysis, the system models the dynamic changes of the skeleton model to obtain the customer's standing posture, gait and posture change patterns. Temporal convolution processing is performed on the joint motion trajectory established based on the 3D skeleton model, and the behavioral features of customers are extracted using a temporal convolutional network; the rate of change of body posture of customers in different time windows is calculated, and the gait cycle, limb swing frequency and head rotation angle are modeled to identify the trend of customers' business needs. Key point detection is performed on the customer's facial area to extract the coordinates of facial feature points, and feature changes under different expression states are calculated based on principal component analysis. Time series analysis is performed on the curvature of the corners of the mouth, the rise and fall of the eyebrows, and the opening and closing angle of the eyelids to calculate the rate of expression change and facial muscle movement characteristics, thereby obtaining the customer's emotional state.
[0012] As a preferred embodiment of the electricity marketing customer behavior identification method based on multimodal analysis described in this invention, the method includes: deploying temperature and humidity sensors, an air quality detection module, and a noise monitoring device to acquire the temperature, humidity, carbon dioxide concentration, air flow rate, and environmental noise level of the business hall, and adding timestamps to the acquired data. Based on the collected temperature and humidity data, the fluctuations in temperature and humidity are analyzed, and the trends of temperature and humidity changes are extracted. The temperature and humidity data are compared with the customer's head movement trajectory, limb movement frequency, and facial expression features extracted from the video data to calculate the customer's gait rate, dwell time, frequency of micro-body movements, and changes in facial muscle activity under different temperature and humidity conditions. When the temperature and humidity exceed the set threshold and the customer's behavior pattern changes abnormally, the behavior confidence of the video data is adjusted to correct the misjudgments caused by temperature and humidity fluctuations and to correct the customer behavior recognition results. Based on air quality data, the impact of carbon dioxide concentration and air flow rate on customer behavior patterns is analyzed; combined with the speech temporal characteristics of speech data extraction methods, the impact of air quality on the speech rate, volume, pitch stability and speech clarity of customer speech signals is analyzed; the degree of interference of air quality on customer speech feature parameters is calculated, and a compensation algorithm is used to correct the speech signals affected by the environment. Based on data from noise monitoring equipment, the noise interference coefficient in the business hall is calculated, and combined with the beamforming processing results of the voice data extraction method, the signal-to-noise ratio change of the voice signal is judged. When the ambient noise exceeds the set threshold, the beam direction of the microphone array is adjusted to enhance the signal strength of the target customer's voice, and the confidence level of the voice data is dynamically adjusted.
[0013] As a preferred embodiment of the electricity marketing customer behavior identification method based on multimodal analysis described in this invention, the feature extraction through the cross-modal feature alignment method includes assigning timestamps to the target customer's voice data and customer video data respectively, and performing synchronization alignment based on a global clock so that the voice data and video data at the same moment have a unified time identifier. Based on synchronized voice and video data, a cross-modal Transformer feature alignment model is constructed to calculate the temporal correspondence between features of different modalities and extract highly correlated features. Based on the time synchronization information of voice data and video data, establish the correspondence between voice expression patterns and body movement patterns; Based on a feature-level fusion network, the emotional features of voice data and the behavioral features of video data are mapped in a high dimension to generate a fused customer behavior vector.
[0014] As a preferred embodiment of the power marketing customer behavior identification method based on multimodal analysis described in this invention, the real-time analysis includes nonlinear feature mapping based on a stacked autoencoder, and extracting features with representation capabilities higher than a set threshold through dimensionality reduction, thereby removing redundant information. A self-attention mechanism is used to calculate the feature weights of the behavior vector at different time steps and extract key behavior patterns. A bidirectional long short-term memory network is used to perform time series modeling on the fused customer behavior vector, calculate the temporal correlation of behavior patterns, and model historical behavior trends to enhance the ability to identify continuous behavior. A cross-modal fusion mechanism is constructed in the output layer of a bidirectional long short-term memory network. The cross-modal information entropy of behavioral patterns is calculated based on a multilayer perceptron, and the current behavioral state of customers is identified through a nonlinear classification method. Deep reinforcement learning is used to optimize model parameters, with the accuracy of customer behavior recognition as the reward signal, and the model hyperparameters are dynamically adjusted through a policy gradient optimization algorithm.
[0015] As a preferred embodiment of the electricity marketing customer behavior identification method based on multimodal analysis described in this invention, the prediction of subsequent customer behavior includes using a time series prediction method to construct a customer behavior time series model and using a gated loop unit to predict the trend of customer behavior changes at different time steps. The system calculates the temporal variation trend of the fused customer behavior vector, uses convolutional temporal networks to analyze customer behavior patterns in different time windows, and predicts whether customers will continue to consult, leave the business hall, or escalate their emotions.
[0016] As a preferred embodiment of the power marketing customer behavior recognition system based on multimodal analysis described in this invention, the data acquisition module is used to deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data of customer voice, video and environment; The feature fusion module is used to perform time synchronization alignment on the collected multimodal data and to extract features using a cross-modal feature alignment method; and... The behavior recognition module is used to input the fused customer behavior vectors into the customer behavior recognition model for real-time analysis, to determine the customer's business needs, predict the customer's subsequent behavior, and optimize business processes based on customer behavior data.
[0017] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the electricity marketing customer behavior identification method based on multimodal analysis as described above.
[0018] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the electricity marketing customer behavior identification method based on multimodal analysis as described above.
[0019] The beneficial effects of this invention: The power marketing customer behavior recognition method based on multimodal analysis provided by this invention first establishes a time synchronization mechanism through cross-modal Transformer feature alignment, ensuring that voice, video, and environmental data are fused at the same time step, thus solving the feature misalignment problem caused by different data time steps in existing technologies. In traditional multimodal analysis methods, the voice sampling rate and video frame rate are different, which can easily lead to information mismatch when directly fused, affecting the accuracy of behavior recognition. Through the Dynamic Time Warping (DTW) algorithm, this invention can automatically adjust the time step of different modal data, so that voice, video, and other information are aligned before fusion, thereby eliminating the behavior judgment bias caused by data asynchrony. The accuracy of time synchronization directly determines the effectiveness of subsequent feature fusion. Therefore, based on this, this invention further adopts a feature-level fusion network to perform high-dimensional mapping of voice emotion and video behavior features, enabling comprehensive judgment of customer emotions and actions, and improving the overall behavior recognition accuracy.
[0020] Building upon accurate behavior recognition and prediction, this invention further optimizes business processes to enhance customer experience and improve branch office operational efficiency. Traditional branch office service processes suffer from long customer wait times and uneven distribution of window resources, and existing technologies struggle to optimize resource allocation based on real-time customer behavior data. This invention employs an intelligent window allocation optimization strategy, combined with reinforcement learning to optimize customer behavior prediction results, automatically adjusting branch office window resource allocation strategies. This prioritizes high-traffic windows for relevant customers, reducing wait times and improving overall operational efficiency. Simultaneously, by combining customers' historical transaction records and real-time behavior status, this invention can proactively push relevant business guidelines based on predictive models, enabling customers to quickly understand the process, increasing the rate of self-service transactions, thereby reducing the workload of human customer service and enhancing the branch office's service capabilities.
[0021] Traditional speech recognition systems often perform poorly in noisy environments. This invention, however, incorporates an environmental data compensation algorithm. During speech recognition, it uses data from noise monitoring equipment to calculate the noise interference coefficient within the sales hall. An adaptive beamforming strategy automatically adjusts the microphone array's pickup direction, enhancing the signal strength of the target customer's voice and improving the signal-to-noise ratio. Furthermore, in environments with high temperature and humidity or poor air quality, customer behavior patterns may change; for example, gait and facial expressions may be affected by the external environment. This invention, combined with an environmental data compensation mechanism, automatically adjusts the confidence level of video data during behavior recognition, enabling the system to maintain high stability in complex environments and improving its applicability. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is an overall flowchart of a method for identifying customer behavior in electricity marketing based on multimodal analysis, provided as an embodiment of the present invention. Detailed Implementation
[0024] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0025] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0026] Example 1, referring to Figure 1 As an embodiment of the present invention, a method for identifying customer behavior in electricity marketing based on multimodal analysis is provided, comprising: Step 101: Deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data of the customer's voice, video, and environment; Step 102: Perform time synchronization alignment on the collected multimodal data, and extract features using a cross-modal feature alignment method; Step 103: Input the fused customer behavior vector into the customer behavior recognition model for real-time analysis, determine the customer's business needs, predict the customer's subsequent behavior, and optimize business processes based on customer behavior data.
[0027] A microphone array is deployed to acquire a mixed audio signal containing the target customer's voice and environmental noise. The acquired audio signal is processed by beamforming. Based on the time difference, phase difference, and amplitude changes of the signals received by different microphones, the direction of the sound source is calculated, and the pickup weight of the microphone array is adjusted to enhance the target customer's voice signal while reducing noise interference from other directions. Ultrasonic sensors are deployed to emit ultrasonic signals and receive echo signals reflected from environmental objects and the human body. Based on the time of flight of the ultrasonic signals, the time difference of echoes at different locations is calculated, and combined with the direction of the sound source identified by beamforming technology, the location of the customer's voice is determined. Frequency modulation analysis is performed on the ultrasonic signals to distinguish between the customer's voice echo and the environmental noise echo, and the positioning data is matched with the pickup direction of the microphone array to optimize the extraction of the target customer's voice signal. Short-time Fourier transform (SFT) is performed on the target customer's speech signal after beamforming and voice location detection to decompose the time-frequency characteristics of the audio signal. The spectrum after the SFT is mapped using a Mel filter bank, and Mel spectral features are extracted. Based on the extracted Mel spectrum, the spectral envelope, frequency band energy distribution, and formant features of the speech signal are calculated to obtain the timbre characteristics of the speech signal. The time-domain characteristics of the speech signal are analyzed by calculating the speech rate, energy fluctuations, pause intervals, and fundamental frequency trajectory to form the speech temporal characteristics. Based on the extracted Mel spectrum features and speech temporal features, an emotion classification model is constructed. The spectral variation trend, speech rate pattern and energy distribution of the speech signal in different time windows are analyzed, the dynamic changes of pitch, intensity and rhythm are calculated, and feature parameters that can characterize emotional state are extracted. The classification model is used to analyze the feature parameters, and combined with the variation trend of temporal features, the customer's emotion category is identified, and the speech emotion recognition result is output. The system associates the identified customer voice signals and their corresponding emotion categories with the customer's voice location data, and interfaces the temporal characteristics of the voice signals with the subsequent customer behavior analysis module to provide the voice data input required for customer behavior pattern analysis.
[0028] Deploy camera equipment to acquire continuous video streams containing customer head movements, gestures, and body movements, and perform frame-by-frame processing on the video data to extract frame sequences for behavioral analysis; Target detection is performed on the video frame sequence, and the positions of the customer's head, hands and torso are extracted using a human pose detection algorithm. The coordinates of the detected key points are smoothed based on the Kalman filter algorithm, and the motion trajectory between consecutive frames is analyzed by combining optical flow method to establish the motion time series of the customer's head, gestures and limb movements. The system uses cameras to acquire the three-dimensional spatial coordinates of various joints on the customer's body, performs three-dimensional modeling of key joints (head, shoulder, elbow, wrist, hip, knee, and ankle, etc.), and uses Euclidean distance to calculate the spatial relationship between different joints to generate a skeletal structure model. Based on time series analysis, the system models the dynamic changes of the skeletal model to obtain the customer's standing posture, gait, and posture change patterns. Temporal convolution processing is performed on the joint motion trajectory established based on the 3D skeleton model, and the behavioral features of customers are extracted using a temporal convolutional network; the rate of change of body posture of customers in different time windows is calculated, and the gait cycle, limb swing frequency and head rotation angle are modeled to identify the trend of customers' business needs. Key point detection is performed on the customer's facial area to extract the coordinates of facial feature points (corner of mouth, eyebrows, eyes, etc.), and feature changes under different expression states are calculated based on principal component analysis. Time series analysis is performed on the curvature of the corner of the mouth, the rise and fall of the eyebrows, and the opening and closing angle of the eyelids to calculate the rate of expression change and facial muscle movement characteristics, thereby obtaining the customer's emotional state. The system associates customer head movements, gestures, body language data with facial expression data to form complete behavioral analysis data. This data is then transmitted to the customer behavior recognition module to provide visual data input for identifying customer business needs.
[0029] Temperature and humidity sensors, air quality detection modules, and noise monitoring equipment are deployed to acquire the temperature, humidity, carbon dioxide concentration, air flow rate, and environmental noise level of the business hall, respectively. Timestamps are added to the acquired data to keep it synchronized with the acquisition time of voice and video data. Based on the collected temperature and humidity data, the fluctuations in temperature and humidity are analyzed, and the trends of temperature and humidity changes are extracted. The temperature and humidity data are compared with the customer's head movement trajectory, limb movement frequency, and facial expression features extracted from the video data to calculate the customer's gait rate, dwell time, frequency of micro-body movements, and changes in facial muscle activity under different temperature and humidity conditions. When the temperature and humidity exceed the set threshold and the customer's behavior pattern changes abnormally, the behavior confidence of the video data is adjusted to correct the misjudgments caused by temperature and humidity fluctuations and to correct the customer behavior recognition results. Based on air quality data, the impact of carbon dioxide concentration and air flow rate on customer behavior patterns is analyzed; combined with the speech temporal features of the speech data extraction method, the impact of air quality on the speech rate, volume, pitch stability and speech clarity of customer speech signals is analyzed; the degree of interference of air quality on customer speech feature parameters is calculated, and a compensation algorithm is used to correct the speech signals affected by the environment in order to improve the accuracy of speech emotion recognition. Based on data from noise monitoring equipment, the noise interference coefficient in the business hall is calculated, and combined with the beamforming processing results of the voice data extraction method, the signal-to-noise ratio change of the voice signal is judged. When the ambient noise exceeds the set threshold, the beam direction of the microphone array is adjusted to enhance the signal strength of the target customer's voice, and the confidence of the voice data is dynamically adjusted to optimize the stability of voice recognition. Based on the results of temperature and humidity impact verification, air quality impact verification, and noise environment impact verification, a comprehensive analysis is conducted on the customer behavior characteristics obtained by the video data extraction method and the voice emotion recognition results of the voice data extraction method. When environmental factors are detected that may cause errors in behavior recognition or voice recognition results, an environmental compensation mechanism is triggered to adjust the behavior patterns or voice emotion states that are prone to misjudgment under specific environmental conditions. The current recognition results are compared with historical behavior patterns to optimize the accuracy and robustness of customer behavior recognition results.
[0030] Feature extraction through cross-modal feature alignment includes assigning timestamps to the target customer's voice data and customer's video data respectively, and performing synchronization alignment based on a global clock so that voice data and video data at the same moment have a unified time identifier. Based on synchronized voice and video data, a cross-modal Transformer feature alignment model is constructed to calculate the temporal correspondence between features of different modalities and extract highly correlated features. Calculate the temporal correlation between speech rate changes, pitch fluctuations, and tone variations in speech data and mouth corner curvature, eyebrow elevation, and eye direction changes in video data; By combining the pause duration, energy fluctuations, and pronunciation clarity of speech data with head micro-movements and neck posture adjustments in video data, a joint model is performed to generate cross-modal alignment vectors. By utilizing adaptive attention weights, the temporal matching effect of speech-video joint features is enhanced, ensuring that speech emotion features and facial expression features are expressed synchronously within the same time window.
[0031] Based on the time synchronization information of voice data and video data, establish the correspondence between voice expression patterns and body movement patterns; Extract speech rhythm, pause intervals, and fundamental frequency trajectory from the speech data, and perform time-series matching with customer limb movement trajectory, gait rhythm, and changes in body center of gravity from the video data; A dynamic time warping method is used to non-linearly align the time-series features of speech data and the motion trajectory of video data, and to calculate the speech-action temporal correlation matrix. By combining gait detection, we can analyze the customer's body balance, standing stability and body movement coordination during the speech process to characterize the synergistic relationship between the customer's speech rhythm and body movements. Based on the fused customer behavior vectors, structured behavior description data is generated and input into the customer behavior recognition model; Predict customer behavior based on multimodal features such as voice, body language, facial expressions, and gait; By combining time-series behavioral data, a customer business needs identification model is established, and personalized business recommendations are optimized. By analyzing customer behavior patterns during voice communication, we can determine the urgency of their business needs and trigger corresponding service response mechanisms to improve customer service efficiency.
[0032] Based on a feature-level fusion network, the emotional features of voice data and the behavioral features of video data are mapped in a high dimension to generate a fused customer behavior vector.
[0033] Emotional features of speech data include speech energy distribution, pitch variation, and tone fluctuation patterns; Behavioral characteristics of video data include changes in facial expressions, eye tracking, micro-head movements, and coordination of bodily movements; Self-supervised learning is used to optimize the speech-video feature embedding method, which makes the speech emotion features and behavior features more separable in the feature space and improves the generalization ability of behavior recognition. An adaptive normalization strategy is adopted to ensure that the feature vectors of different modal data are calculated on a uniform scale, making the distribution of the fused data more stable and improving the fusion effect of cross-modal features.
[0034] Real-time analysis includes nonlinear feature mapping based on stacked autoencoders, and extracting features with representation capabilities higher than a set threshold through dimensionality reduction, removing redundant information, and improving the effectiveness of feature representation. A self-attention mechanism is employed to calculate the feature weights of the behavior vector at different time steps and extract key behavior patterns to enhance the modeling ability for short-term and long-term dependent behaviors. A bidirectional long short-term memory network is used to perform time series modeling on the fused customer behavior vector, calculate the temporal correlation of behavior patterns, and model historical behavior trends to enhance the ability to identify continuous behavior. A cross-modal fusion mechanism is constructed in the output layer of a bidirectional long short-term memory network. The cross-modal information entropy of behavioral patterns is calculated based on a multilayer perceptron, and the current behavioral state of customers is identified through a nonlinear classification method. Deep reinforcement learning is used to optimize model parameters. The accuracy of customer behavior recognition is used as a reward signal. The model hyperparameters are dynamically adjusted through a policy gradient optimization algorithm to improve the generalization ability of customer behavior recognition.
[0035] Predicting subsequent customer behavior involves using time series forecasting methods to construct a time series model of customer behavior and using gated loop units to predict the trend of customer behavior changes at different time steps. The temporal variation trend of the fused customer behavior vector is calculated, and the behavior pattern of customers in different time windows is analyzed using convolutional temporal networks to predict whether customers will continue to consult, leave the business hall, or escalate their emotions. By employing a multi-task learning framework, the model simultaneously predicts customer business demand categories and subsequent behavioral trends, thereby improving the overall generalization ability of the prediction model.
[0036] Example 2, an embodiment of the present invention, provides a power marketing customer behavior identification system based on multimodal analysis, comprising: The data acquisition module is used to deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data of the customer's voice, video and environment; The feature fusion module is used to perform time synchronization alignment on the collected multimodal data and to extract features using a cross-modal feature alignment method; and... The behavior recognition module is used to input the fused customer behavior vectors into the customer behavior recognition model for real-time analysis, to determine the customer's business needs, predict the customer's subsequent behavior, and optimize business processes based on customer behavior data.
[0037] Example 3, an embodiment of the present invention, differs from the previous two embodiments in that: If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0038] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0039] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0040] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0041] Example 4, an embodiment of the present invention, provides a method for identifying customer behavior in electricity marketing based on multimodal analysis, comprising: Step 1: Deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data of the customer's voice, video, and environment; 1.1 For language data: Deploy a microphone array to acquire a mixed audio signal containing the target customer's voice and environmental noise; perform beamforming processing on the acquired audio signal, calculate the direction of the sound source based on the time difference, phase difference, and amplitude changes of the signals received by different microphones, and adjust the pickup weight of the microphone array to enhance the target customer's voice signal while reducing noise interference from other directions. Ultrasonic sensors are deployed to emit ultrasonic signals and receive echo signals reflected from environmental objects and the human body. Based on the time of flight of the ultrasonic signals, the time difference of echoes at different locations is calculated, and combined with the direction of the sound source identified by beamforming technology, the location of the customer's voice is determined. Frequency modulation analysis is performed on the ultrasonic signals to distinguish between the customer's voice echo and the environmental noise echo, and the positioning data is matched with the pickup direction of the microphone array to optimize the extraction of the target customer's voice signal. After beamforming and voice location detection, the target customer's speech signal undergoes a short-time Fourier transform to decompose the time-frequency characteristics of the audio signal. The spectrum after the short-time Fourier transform is then energy-mapped using a Mel filter bank, and Mel spectral features are extracted. Based on the extracted Mel spectrum, the spectral envelope, frequency band energy distribution, and formant features of the speech signal are calculated to obtain the timbre characteristics of the speech signal. Finally, the temporal characteristics of the speech signal are analyzed by calculating the speech rate, energy fluctuations, pause intervals, and fundamental frequency trajectory to form the speech temporal characteristics. Based on the extracted Mel spectrum features and speech temporal features, an emotion classification model is constructed. The spectral variation trend, speech rate pattern and energy distribution of the speech signal in different time windows are analyzed, the dynamic changes of pitch, intensity and rhythm are calculated, and feature parameters that can characterize emotional state are extracted. The classification model is used to analyze the feature parameters, and combined with the variation trend of temporal features, the customer's emotion category is identified, and the speech emotion recognition result is output. The system associates the identified customer voice signals and their corresponding emotion categories with the customer's voice location data, and interfaces the temporal characteristics of the voice signals with the subsequent customer behavior analysis module to provide the voice data input required for customer behavior pattern analysis.
[0042] A microphone array is deployed to acquire a mixed audio signal containing the target customer's voice and ambient noise. The microphone array consists of M microphones and is used to collect the target customer's voice signal. Array processing technology is then used to enhance the target voice while suppressing ambient noise.
[0043] An ultrasonic sensor is deployed to emit ultrasonic signals and receive echo signals reflected from environmental objects and the human body to measure the target customer's vocal location. The time of flight of the ultrasonic signal is calculated from the propagation path distance from the customer to the ultrasonic sensor and the speed of sound in the air. This vocal location information is used to optimize the beam direction of the microphone array, improving the accuracy of target speech pickup while reducing interference signals from other directions.
[0044] The acquired audio signals are processed using beamforming. The direction of the sound source is calculated based on the time difference, phase difference, and amplitude changes of the signals received by different microphones, and the pickup weights of the microphone array are adjusted. The azimuth and elevation angles of the target speech are also determined. The optimized beam direction enhances the signal strength of the target customer's voice and reduces interference from non-target voice and background noise.
[0045] Short-time Fourier Transform (STFT) is performed on the target customer's speech signal after beamforming and voice location detection to decompose the time-frequency characteristics of the audio signal. Frequency domain information of the speech signal is then extracted using STFT for subsequent feature analysis.
[0046] The spectrum after STFT transformation is energy mapped using the Mel filter bank, and Mel spectral features are extracted.
[0047] The extracted Mel spectrum features are used for voice emotion analysis to identify the customer's emotional state.
[0048] The temporal characteristics of speech signals are analyzed by calculating speech rate, energy fluctuations, pause intervals, and fundamental frequency trajectory to form speech temporal features. Speech rate is calculated based on the time variation of syllables, obtaining the time interval between syllables through short-time energy detection. Energy fluctuations are calculated based on the amplitude of short-time energy fluctuations to identify the intensity of emotional expression in the speech. Pause intervals are calculated based on the duration of silent segments with short-time energy below a threshold to measure the customer's thinking time. The fundamental frequency trajectory is obtained through autocorrelation or harmonic peak detection methods, analyzing the rising or falling trends of pitch to reflect the emotional characteristics of the speech expression.
[0049] Based on extracted Mel-frequency spectral features and speech temporal features, an emotion classification model is constructed. The spectral variation trend, speech rate pattern, and energy distribution of the speech signal within different time windows are analyzed, and the dynamic changes in pitch, intensity, and rhythm are calculated. Pitch is extracted and calculated from fundamental frequency features, representing the tonic pitch of the speech; intensity is calculated from short-time energy, representing the loudness of the speech; and rhythm is analyzed through dynamic time warping to determine the changes in syllable intervals. Finally, a self-attention mechanism is used to calculate the temporal weights of the speech emotion features, and a deep learning classification model is employed for emotion recognition. The classification model calculates the probability of the emotion category based on Softmax, and the output of the classification model is the customer's emotion category to support subsequent customer behavior analysis and service strategy optimization.
[0050] 1. Speech signal acquisition: 1.1 Setting up the microphone array: Deploy a microphone array to acquire a mixed audio signal containing the target customer's voice and ambient noise. Assume the microphone array consists of... It consists of 1 microphone, and its spatial location is as follows: ; in, For microphone Position vector in three-dimensional space These represent the three-dimensional coordinates of the microphone. This represents the total number of microphones.
[0051] 1.2 Deployment of ultrasonic sensors: Ultrasonic sensors are deployed to emit ultrasonic signals and receive echo signals reflected from environmental objects and the human body to measure the location of the target customer's voice.
[0052] 2. Target speech enhancement: 2.1 Calculate the direction of the sound source: Calculate the time difference of reception (TDOA) of the microphone array. Phase difference ) and amplitude difference ( Based on this, the azimuth angle of the customer's voice is estimated. and elevation angle : ; in, The speed of sound in air. The time difference between signals received by different microphone channels. The difference in signal amplitude received by different microphones, For microphone and The spatial distance between them.
[0053] 2.2 Optimize beam direction by combining ultrasonic measurement: Calculate the customer's vocal location based on the time-of-flight of ultrasonic signals: ; in, The propagation path distance from the customer to the ultrasonic sensor. This represents the speed at which ultrasound travels through the air.
[0054] Optimize beamforming direction using ultrasonic positioning results: ; ; in, The noise covariance matrix is... For the first The signal vector received by each microphone The average noise signal received by the microphone. To combine the customer's sound source direction vector measured by ultrasound, This represents the beamforming weight vector. This indicates transpose.
[0055] 3. Speech Feature Extraction: 3.1 Calculate the Short-Time Fourier Transform (STFT): Perform a short-time Fourier transform (STFT) on the enhanced target speech signal: ; in, For time frequency The Fourier transform result at the point, For input voice signal, For time window functions, Represents a time frame. Indicates frequency, This represents the complex exponential basis functions of the Fourier transform.
[0056] 3.2 Extracting Mel spectral features: Calculate the energy distribution of the Mel spectrum: ; in, Let Mel be the weighting function of the filter. is the number of filter banks, and k represents the index number of the filter bank.
[0057] Calculate the spectral envelope, frequency band energy distribution, and formant characteristics of a speech signal: ; 3.3 Speech temporal feature analysis: Calculate the speech rate, energy fluctuations, pause intervals, and fundamental frequency trajectory of the speech signal: ; in, For the short-time energy entropy of speech, For the first The short-time energy of a frame This represents the short-time energy normalization probability.
[0058] 4. Voice Emotion Classification: 4.1 Calculate emotion-related parameters: Calculate the dynamic changes in pitch, energy, and rhythm: ; in, For the first The fundamental frequency of the frame, For the first The short-time energy of a frame The time interval between adjacent frames. This represents the energy difference between adjacent frames.
[0059] 4.2 Identifying emotion categories using classification models: Based on the extracted Mel spectrum features and speech temporal features, an emotion classification model is constructed: ; in, This represents the attention matrix of the hidden layer of the neural network. Represents the query matrix. Represents the key matrix, Represents a value matrix, This represents the scaling factor for the feature dimension.
[0060] Final emotion classification output: ; in, This represents the output of the final classification prediction. This represents the activation function. This represents the weight matrix of the neural network. This represents the attention feature matrix calculated above. This represents the bias vector.
[0061] 1.2 For video data: Deploy camera equipment to acquire continuous video streams containing customer head, gestures, and body movements, and perform frame segmentation on the video data to extract frame sequences for behavioral analysis; Target detection is performed on the video frame sequence, and the positions of the customer's head, hands and torso are extracted using a human pose detection algorithm. The coordinates of the detected key points are smoothed based on the Kalman filter algorithm, and the motion trajectory between consecutive frames is analyzed by combining optical flow method to establish the motion time series of the customer's head, gestures and limb movements. The system uses cameras to acquire the three-dimensional spatial coordinates of various joints in the customer's body, performs three-dimensional modeling of the joints (key joints such as head, shoulder, elbow, wrist, hip, knee and ankle), and uses Euclidean distance to calculate the spatial relationship between different joints to generate a skeletal structure model. Based on time series analysis, the system models the dynamic changes of the skeletal model to obtain the customer's standing posture, gait and posture change patterns. Temporal convolution processing is performed on the joint motion trajectory established based on the 3D skeleton model, and the behavioral features of customers are extracted using a temporal convolutional network; the rate of change of body posture of customers in different time windows is calculated, and the gait cycle, limb swing frequency and head rotation angle are modeled to identify the trend of customers' business needs. Key point detection is performed on the customer's facial area to extract the coordinates of facial feature points (corner of mouth, eyebrows, eyes, etc.), and feature changes under different expression states are calculated based on principal component analysis. Time series analysis is performed on the curvature of the corner of the mouth, the rise and fall of the eyebrows, and the opening and closing angle of the eyelids to calculate the rate of expression change and facial muscle movement characteristics, thereby obtaining the customer's emotional state. The system associates customer head movements, gestures, body language data with facial expression data to form complete behavioral analysis data. This data is then transmitted to the customer behavior recognition module to provide visual data input for identifying customer business needs.
[0062] 1.2.1 Video Stream Acquisition and Preprocessing: Step Description: In this step, the first step is to deploy video surveillance equipment within the sales area and record high-definition video of customer activities in different areas. The deployed video surveillance equipment should have wide-angle shooting capabilities to cover the entire sales area and support autofocus and lighting compensation to ensure clear image data under various lighting conditions. Furthermore, the video surveillance equipment must have a high frame rate (at least 30 FPS) to ensure that subsequent behavioral analysis can accurately capture customers' body movements and micro-expression changes.
[0063] The acquired video stream data may have issues such as lighting changes, camera shake, and background interference, therefore preprocessing is required to optimize data quality: 1. Gaussian filter noise reduction: used to remove sensor noise and improve video quality; 2. Color normalization: Histogram equalization is used to enhance visual contrast under different lighting conditions; 3. Frame serialization: Convert video data into time-series frames to support subsequent behavior analysis.
[0064] Algorithm Implementation: Let time be... The captured video frame sequence is as follows: ; in, Indicates time The captured video frame sequence, Indicates the first Frame image, Indicates the length of the frame sequence.
[0065] To remove noise, a Gaussian filter is used: ; in, For the denoised first frame, It is a two-dimensional Gaussian filter kernel.
[0066] The denoised frame data will be used for target detection and key point extraction.
[0067] 1.2.2 Target Detection and Keypoint Extraction: Step Description: The main objective of this step is to extract key points of the customer's head, gestures, and limbs from the preprocessed video data to build the foundational information for customer behavior analysis. This process consists of two steps: 1. Object Detection: YOLOv5 is used for human detection to determine the bounding box of the customer's location; 2. Keypoint Detection: Using OpenPose to identify key points of the customer's head, hands, and limbs to build a complete human skeleton model.
[0068] Algorithm implementation: The bounding box representation of YOLOv5 object detection is as follows: ; in, Indicates time The bounding box set of the target customer The coordinates of the top left corner of the bounding box. These are the coordinates of the bottom right corner of the bounding box.
[0069] Joint coordinates obtained from keypoint detection: ; in, For a moment The set of key points to be identified Key point Coordinates in a two-dimensional image For confidence level, Indicates the key point number. This represents the set of key points on the human body.
[0070] The extracted key points will be used for motion trajectory calculation to analyze customer behavior patterns.
[0071] 1.2.3 Calculation of motion trajectory: Step Description: The core of this step is calculating the customer's motion trajectory in the video to characterize their behavioral patterns. Key challenges include: 1. Video data may contain inter-frame jitter, resulting in unstable trajectory; 2. The customer's movement trajectory may be partially obscured, requiring trajectory completion; 3. It is necessary to distinguish between the customer's active movement and passive movement (such as displacement caused by camera shake).
[0072] This step uses Kalman filtering for trajectory smoothing and combines it with optical flow to calculate trajectory offset, thereby obtaining a highly stable motion trajectory.
[0073] Algorithm implementation: 1.2.3.1 Trajectory Smoothing: Kalman filter predicts the trajectory: ; in, To predict the location of key points, Here is the state transition matrix. To control the input matrix, For noise terms, This indicates the control input (used to correct the direction of motion). Indicates time The key point locations are derived from the key point detection results.
[0074] 1.2.3.2 Calculation of motion trajectory: Calculating motion vectors using optical flow: ;
[0075] in, Key points of optical flow calculation Motion vector, Indicates time and Processed image frames, Indicate key points The set of neighboring pixels, This represents the neighborhood weights (Gaussian distribution).
[0076] Finally, the smooth trajectory and the motion trajectory are merged: ; in, This represents the trajectory after Kalman filtering. This indicates the fusion weight.
[0077] The results will be used for 3D modeling of joints.
[0078] Step description: Calculate the three-dimensional coordinates of the client's limb joints to form a complete skeletal model; Analyze the changes in spatial distance between different key points to infer customer behavior trends.
[0079] Algorithm Implementation: Calculation of 3D Joint Coordinates ; in, Indicates time Key points Coordinates in three-dimensional space Indicate key points Pixel coordinates on the image plane Indicate key points The depth information is calculated as follows: ; in, Indicates the focal length of the camera. This represents the actual physical distance between two key points. It represents the distance between two joints on the image pixel plane.
[0080] Euclidean distance between joints: ; in, Indicates time Key points and The three-dimensional Euclidean distance between them Indicates time Key points The three-dimensional coordinates Indicates time Key points The three-dimensional coordinates.
[0081] 1.2.5 Time Series Behavior Analysis: Step description: This step calculates gait stability and limb swing frequency of customer behavior based on time series analysis.
[0082] Algorithm Implementation: Gait Variation: ; in, Indicates time Gait stability characteristics, This represents the total number of key points detected. Indicates time Key points and The three-dimensional Euclidean distance between them Indicates time Key points and The three-dimensional Euclidean distance between them represents the Euclidean distance at the previous time step.
[0083] Limb swing frequency: ; in, Indicates time The frequency of limb swinging This represents the total number of key points detected. Indicate key points At any moment The velocity vector of motion is calculated as follows: ; in, Indicates time Key points The three-dimensional coordinates Indicates time Key points The three-dimensional coordinates This indicates the time interval between adjacent frames.
[0084] 1.2.6 Multimodal Behavior Fusion: Step Description: This step integrates gait features, facial expression features, and motion information, and uses a multilayer perceptron (MLP) for final feature fusion.
[0085] Algorithm Implementation: Gait Features: ; in, Represents the gait feature vector. This represents the gait feature weight matrix, indicating the weight of gait stability in the final behavior prediction. Indicates time Gait stability characteristics, This represents the gait feature bias term.
[0086] Characteristics of facial expression changes: ; in, This represents the feature vector of facial expression changes. This represents the facial expression feature weight matrix, indicating the weight of facial expressions in the final behavior prediction. Indicates time Information about changes in facial expressions This indicates the bias term for facial expression features.
[0087] Final fusion behavior vector: ; in, This represents the final fusion behavior vector. This represents the feature concatenation operation. MLP stands for Multilayer Perceptron, used for feature transformation and mapping. For a moment The frequency of limb swinging.
[0088] Behavioral classification prediction: ; in, Indicates time Predicted probability distribution of customer behavior categories This represents the classification weight matrix, indicating the contribution of each behavioral feature to the final classification. This represents the classification bias term, and Softmax represents the normalization function, which outputs the customer's behavior category (such as standing, sitting, leaving, emotional changes, etc.).
[0089] 1.3 For environmental data: Deploy temperature and humidity sensors, air quality detection modules and noise monitoring equipment to acquire the temperature, humidity, carbon dioxide concentration, air flow rate and environmental noise level of the business hall, and add timestamps to the acquired data to keep it synchronized with the acquisition time of voice data and video data; Based on the collected temperature and humidity data, the fluctuations in temperature and humidity are analyzed, and the trends of temperature and humidity changes are extracted. The temperature and humidity data are compared with the customer's head movement trajectory, limb movement frequency, and facial expression features extracted from the video data to calculate the customer's gait rate, dwell time, frequency of micro-body movements, and changes in facial muscle activity under different temperature and humidity conditions. When the temperature and humidity exceed the set threshold and the customer's behavior pattern changes abnormally, the behavior confidence of the video data is adjusted to correct the misjudgments caused by temperature and humidity fluctuations and to correct the customer behavior recognition results. Based on air quality data, the impact of carbon dioxide concentration and air flow rate on customer behavior patterns is analyzed; combined with the speech temporal features of the speech data extraction method, the impact of air quality on the speech rate, volume, pitch stability and speech clarity of customer speech signals is analyzed; the degree of interference of air quality on customer speech feature parameters is calculated, and a compensation algorithm is used to correct the speech signals affected by the environment in order to improve the accuracy of speech emotion recognition. Based on data from noise monitoring equipment, the noise interference coefficient in the business hall is calculated, and combined with the beamforming processing results of the voice data extraction method, the signal-to-noise ratio change of the voice signal is judged. When the ambient noise exceeds the set threshold, the beam direction of the microphone array is adjusted to enhance the signal strength of the target customer's voice, and the confidence of the voice data is dynamically adjusted to optimize the stability of voice recognition. Based on the results of temperature and humidity impact verification, air quality impact verification, and noise environment impact verification, a comprehensive analysis is conducted on the customer behavior features obtained by the video data extraction method and the voice emotion recognition results of the voice data extraction method. When environmental factors are detected that may cause errors in behavior recognition or voice recognition results, an environmental compensation mechanism is triggered to adjust the behavior patterns or voice emotion states that are prone to misjudgment under specific environmental conditions. The current recognition results are compared with historical behavior patterns. The comparison results are shown in Table 1 to optimize the accuracy and robustness of customer behavior recognition results.
[0090] Table 1. Comparison of Recognition Results Technical Module Limitations of traditional methods This invention is a technical optimization. Improvement effect Voice data collection Unable to accurately extract customer voice in noisy environments Microphone array + beamforming + ultrasonic positioning Improve the quality of the target speech signal and reduce environmental interference. Speech feature extraction Based solely on STFT, lacking in-depth speech analysis Mel spectrum and time-domain feature analysis More accurate identification of customers' emotional state Video Behavior Analysis Limited to 2D keypoint detection, ignoring time dependence. 3D skeleton modeling + time series behavior analysis Improve the continuity and accuracy of customer behavior pattern analysis Facial Recognition Static facial expression analysis, without considering changes over time. Time series analysis + principal component analysis (PCA) More accurately depict changes in customer facial expressions and improve the reliability of emotion recognition. Environmental compensation Ignoring environmental factors leads to misjudgment Temperature and humidity compensation, air quality correction, noise interference adjustment Improve the robustness of voice and behavioral data and reduce environmental errors. Step 2: Perform time synchronization alignment on the collected multimodal data, and extract features using a cross-modal feature alignment method; The target customer's voice data and customer's video data are assigned timestamps respectively, and synchronized and aligned based on a global clock, so that voice data and video data at the same moment have a unified time identifier. Based on synchronized speech and video data, a cross-modal Transformer feature alignment model is constructed to calculate the temporal correspondence between features of different modalities and extract highly correlated features; the temporal correlation between speech rate changes, pitch fluctuations, and tone variations in speech data and mouth corner curvature, eyebrow raising amplitude, and eye direction changes in video data is calculated. By combining the pause duration, energy fluctuations, and pronunciation clarity of speech data with head micro-movements and neck posture adjustments in video data, a joint model is performed to generate cross-modal alignment vectors. By utilizing adaptive attention weights, the temporal matching effect of speech-video joint features is enhanced, ensuring that speech emotion features and facial expression features are expressed synchronously within the same time window; Based on the time synchronization information of voice data and video data, establish the correspondence between voice expression patterns and body movement patterns; Extract speech rhythm, pause intervals, and fundamental frequency trajectory from the speech data, and perform time-series matching with customer limb movement trajectory, gait rhythm, and changes in body center of gravity from the video data; A dynamic time warping method is used to non-linearly align the time-series features of speech data and the motion trajectory of video data, and to calculate the speech-action temporal correlation matrix. By combining gait detection, we can analyze the customer's body balance, standing stability and body movement coordination during the speech process to characterize the synergistic relationship between the customer's speech rhythm and body movements. Based on the fused customer behavior vectors, structured behavior description data is generated and input into the customer behavior recognition model; Predict customer behavior based on multimodal features such as voice, body language, facial expressions, and gait; By combining time-series behavioral data, a customer business needs identification model is established, and personalized business recommendations are optimized. By analyzing customer behavior patterns during voice communication, we can determine the urgency of their business needs and trigger corresponding service response mechanisms to improve customer service efficiency.
[0091] Based on a feature-level fusion network, the emotional features of speech data and the behavioral features of video data are mapped in a high dimension to generate a fused customer behavior vector. The emotional features of speech data include speech energy distribution, pitch variation, and pitch fluctuation pattern. Behavioral characteristics of video data include changes in facial expressions, eye tracking, micro-head movements, and coordination of bodily movements; Self-supervised learning is used to optimize the speech-video feature embedding method, which makes the speech emotion features and behavior features more separable in the feature space and improves the generalization ability of behavior recognition. An adaptive normalization strategy is adopted to ensure that the feature vectors of different modal data are calculated on a uniform scale, making the distribution of the fused data more stable and improving the fusion effect of cross-modal features.
[0092] Because speech and video features have different sampling rates, we need to adjust their timelines to ensure data at the same time step can be matched. To this end, we use the Dynamic Time Warping (DTW) method to calculate the speech time step. and video time step Alignment error between them ensures that data from different modalities have the best matching relationship at the same time step.
[0093] 2.1.1 Calculate the time step error: ; in, Indicates time step Matching error at the location, Indicates time step The speech temporal features at that location, Indicates time step Video behavioral characteristics at the location, Indicates only along the time axis Change (increased speech step size) Indicates only along the time axis Change (video step size increased) This indicates that two time steps are adjusted synchronously (alignment changes).
[0094] 2.1.2 Calculate the optimal alignment path: Calculate the optimal path based on minimum cost search (Dynamic Programming, DP): ; Where WarpingPath represents the error matrix calculated in DTW. In the sequence of time step indices matched by the minimum cost path, This represents the index of the optimal matching time step, calculated by dynamic programming. This represents the aligned speech features. This represents the aligned video features.
[0095] Calculation method: The backtracking algorithm is used to trace back from the final path point to find the optimal time mapping path, and the original feature sequence is interpolated and adjusted.
[0096] 2.2 Cross-modal Transformer feature alignment: After time-step alignment, we further employ a cross-modal Transformer to compute speech features. and video features The temporal correlation ensures that the behavioral features of the two modalities remain semantically consistent.
[0097] 2.2.1 Calculate the cross-modal attention matrix ; in, This represents the cross-modal attention matrix, which measures the degree of matching between speech and video features at different time steps. A query vector representing speech features. Key vectors representing video features Denotes the linear matrix used for eigenvalue transformation. This represents the feature dimension scaling factor.
[0098] 2.2.2 Calculate the cross-modal alignment vector: ; in, The cross-modal projection matrix is learned by minimizing the semantic alignment loss. This indicates the alignment features after merging.
[0099] To find the optimal projection matrix Minimize semantic alignment loss : ; in, This represents the semantic alignment loss, which measures speech features. and aligned video features The Euclidean distance between them ensures that their features are as similar as possible at the same time step. Indicates time step The speech features at that location were obtained after dynamic time warping. Indicates time step Aligned video features at the location.
[0100] Optimization via gradient descent Make Minimum: ; in, Indicates the first The projection matrix of the round iteration, The learning rate controls the optimization step size. This represents the gradient of the loss function with respect to the projection matrix, controlling the update direction.
[0101] The optimization process eventually converges, resulting in the cross-modal projection matrix. It has optimal semantic alignment capabilities.
[0102] 2.3 Voice and Video Behavior Matching Analysis: After aligning the voice and video features, we further calculate the correlation between voice pauses, energy changes, and video motion changes to more accurately describe customer behavior.
[0103] 2.3.1 Calculate changes in speech pauses: ; in, Indicates the amount of change in speech pauses. Indicates time step The speech pause characteristics at that point Indicates time step The speech pause characteristics at the location.
[0104] 2.3.2 Calculate video motion changes: ; in, This indicates the amount of change in the amplitude of video motion. Indicates time step Range of motion of the limbs at that location Indicates time step Range of motion of the limbs at that location.
[0105] 2.3.3 Calculate the matching degree: ; in, This indicates the degree of matching between voice pauses and video motion (the higher the matching degree, the stronger the coordination between voice and motion).
[0106] 2.4 Generate a unified behavior description vector: Finally, by combining the voice and video matching degree and facial expression change features, a complete behavioral description vector is generated.
[0107] 2.4.1 Calculate the matching weights: ; in, The weights representing the degree of audio-video matching. Indicates the characteristics of facial expression changes. This indicates the impact of the control matching degree on the behavioral description.
[0108] 2.4.2 Calculate the final behavioral characteristics: ; in, This represents the final behavior description vector. This represents the weights of cross-modal alignment, action, and facial expression features.
[0109] Final output: ; in, As input for subsequent behavior recognition.
[0110] Step 3: Input the fused customer behavior vector into the customer behavior recognition model for real-time analysis, determine the customer's business needs, predict the customer's subsequent behavior, and optimize business processes based on customer behavior data.
[0111] Nonlinear feature mapping is performed based on stacked autoencoders, and features with representation capabilities exceeding a set threshold are extracted by dimensionality reduction, removing redundant information and improving the effectiveness of feature representation. A self-attention mechanism is employed to calculate the feature weights of the behavior vector at different time steps and extract key behavior patterns to enhance the modeling ability for short-term and long-term dependent behaviors. A bidirectional long short-term memory network is used to perform time series modeling on the fused customer behavior vector, calculate the temporal correlation of behavior patterns, and model historical behavior trends to enhance the ability to identify continuous behavior. A cross-modal fusion mechanism is constructed in the output layer of a bidirectional long short-term memory network. The cross-modal information entropy of behavioral patterns is calculated based on a multilayer perceptron, and the current behavioral state of customers is identified through a nonlinear classification method. Deep reinforcement learning is used to optimize model parameters. The accuracy of customer behavior recognition is used as a reward signal. The model hyperparameters are dynamically adjusted through a policy gradient optimization algorithm to improve the generalization ability of customer behavior recognition.
[0112] A time series forecasting method is used to construct a customer behavior time series model, and a gated loop unit is used to predict the trend of customer behavior changes at different time steps. The temporal variation trend of the fused customer behavior vector is calculated, and the behavior pattern of customers in different time windows is analyzed using convolutional temporal networks to predict whether customers will continue to consult, leave the business hall, or escalate their emotions. By employing a multi-task learning framework, the model simultaneously predicts customer business demand categories and subsequent behavioral trends, thereby improving the overall generalization ability of the prediction model.
[0113] Step 3.1: Dimensionality reduction of behavioral features: The cross-modal alignment behavior vector calculated in step 2 The data has high dimensionality and contains a large amount of redundant information. Directly using this data for time series modeling increases computational complexity and affects prediction accuracy. Therefore, this step uses a stacked autoencoder (SAE) to reduce the dimensionality of behavioral features, extracting the most representative features, removing irrelevant information, and improving data processing efficiency. First, the encoder maps the input features to a lower-dimensional space: ; Then, the input is reconstructed using the decoder to minimize the reconstruction error: ; in, The behavioral characteristics after dimensionality reduction These are the weight matrices for the encoder and decoder, respectively. These are the bias terms for the encoder and decoder, respectively. The activation function is ReLU.
[0114] To ensure that the reduced features still retain important information from the original data, the goal of this step is to minimize the reconstruction error: ; Thus, the behavioral characteristics after dimensionality reduction were obtained. This is for use in subsequent steps.
[0115] Step 3.2: Calculate the temporal behavior weights: In order to capture the weight distribution of customer behavior at different times and to effectively model both short-term and long-term behavioral patterns, this step uses a self-attention mechanism to calculate the weight of behavioral features at different time steps.
[0116] ; in, For time steps The results of behavioral characteristics weighting This represents the query vector at the current time step. This represents the key vector of the previous time step. This represents the value vector at the current time step. As a scaling factor, to prevent excessively large gradients from causing training instability, Softmax calculates attention weights to ensure that the sum of the weights is 1.
[0117] This method can effectively extract key behavioral features from time series data, thereby improving the accuracy of subsequent behavior predictions.
[0118] Step 3.3: Bidirectional LSTM time series modeling: To further explore the temporal correlation of customer behavior, a Bidirectional Long Short-Term Memory (BiLSTM) network is used for time series modeling. This step involves calculating the behavioral feature weights in step 3.2. Based on this, combined with historical behavioral states Calculate the behavior state at the current time step: ; in, For the current time step The behavioral state.
[0119] BiLSTM improves the ability to model behavioral sequences in a temporal manner through bidirectional information transmission.
[0120] This method can learn the trends in customer behavior over time, providing more reliable feature inputs for subsequent behavior prediction.
[0121] Step 3.4: Calculate cross-modal information summary: Because the behavioral patterns of different customers at different times are uncertain, this step calculates the information entropy of cross-modal behavioral features to quantify the uncertainty of behavior in order to optimize the prediction model: ; in, Reflecting the uncertainty of behavioral patterns, This represents the probability distribution of customer behavior categories.
[0122] The optimization goal of this step is: Low entropy ( Small: Stable behavioral patterns that can be directly used for prediction.
[0123] High entropy ( Large: The behavioral patterns are uncertain and require reinforcement learning to optimize.
[0124] Step 3.5: Reinforcement Learning Optimization: When customer behavior is highly uncertain, direct prediction may lead to significant errors. Therefore, this step introduces reinforcement learning policy gradient optimization to improve the customer behavior prediction model, making it adaptable to dynamically changing business environments.
[0125] ; in, To enhance the learning strategy parameters, To optimize the objective function, it is defined as the accuracy of behavior prediction. To optimize the learning rate, control the step size. For gradient update direction, This represents the reinforcement learning policy parameters at time step t. This represents the policy parameters at time step t+1.
[0126] This step involves dynamically adjusting hyperparameters to enable the model to adapt to different customer behavior patterns and improve generalization ability.
[0127] Step 3.6: Predicting Future Behavior: Based on the behavioral characteristics optimized by reinforcement learning, this step predicts the trend of customer behavior changes in future time steps, ensuring intelligent optimization of business processes.
[0128] Time series forecasting using gated recurrent units (GRUs): ; in, Indicates future time steps Customer behavior prediction status.
[0129] Simultaneously, calculate the probability of the behavioral business category: ; in, A weight matrix representing the categories of business behaviors. Indicates the customer's behavior status at a future time step. This represents the bias term in the business forecast.
[0130] Ultimately, predict the customer's future behavior: whether to continue consulting; whether to leave the sales office; whether to escalate their emotions.
[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for identifying customer behavior in electricity marketing based on multimodal analysis, characterized in that, include: Deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data on the customer's voice, video, and environment; The collected multimodal data is time-synchronized and aligned, and features are extracted using a cross-modal feature alignment method. The fused customer behavior vectors are input into the customer behavior recognition model for real-time analysis to determine customer business needs, predict subsequent customer behavior, and optimize business processes based on customer behavior data.
2. The method for identifying customer behavior in electricity marketing based on multimodal analysis as described in claim 1, characterized in that: A microphone array is deployed to acquire a mixed audio signal containing the target customer's voice and environmental noise; beamforming processing is performed on the acquired audio signal, and the direction of the sound source is calculated based on the time difference, phase difference, and amplitude changes of the signals received by different microphones, and the pickup weight of the microphone array is adjusted. Ultrasonic sensors are deployed to emit ultrasonic signals and receive echo signals reflected from environmental objects and the human body. Based on the time of flight of the ultrasonic signals, the time difference of echoes at different locations is calculated, and the direction of the sound source identified by beamforming technology is used to determine the location of the customer's voice. Frequency modulation analysis is performed on the ultrasonic signals to distinguish between the customer's voice echo and the environmental noise echo, and the positioning data is matched with the pickup direction of the microphone array. Short-time Fourier transform is performed on the target customer's speech signal after beamforming and sound location detection to decompose the time-frequency characteristics of the audio signal; the spectrum after short-time Fourier transform is energy-mapped using the Mel filter bank, and Mel spectral features are extracted; based on the extracted Mel spectrum, the spectral envelope, frequency band energy distribution, and formant features of the speech signal are calculated to obtain the timbre features of the speech signal. The temporal characteristics of speech signals are analyzed by calculating the speech rate, energy fluctuations, pause intervals, and fundamental frequency trajectory to form speech temporal features. Based on the extracted Mel spectrum features and speech temporal features, an emotion classification model is constructed. The spectral variation trend, speech rate pattern and energy distribution of the speech signal in different time windows are analyzed, the dynamic changes of pitch, intensity and rhythm are calculated, and feature parameters that can characterize emotional state are extracted. The classification model is used to analyze the feature parameters, and combined with the variation trend of temporal features, the customer's emotion category is identified, and the speech emotion recognition result is output.
3. The method for identifying customer behavior in electricity marketing based on multimodal analysis as described in claim 2, characterized in that: Deploy camera equipment to acquire continuous video streams containing customer head movements, gestures, and body movements, and perform frame-by-frame processing on the video data to extract frame sequences for behavioral analysis; Target detection is performed on the video frame sequence, and the positions of the customer's head, hands and torso are extracted using a human pose detection algorithm. The coordinates of the detected key points are smoothed based on the Kalman filter algorithm, and the motion trajectory between consecutive frames is analyzed by combining optical flow method to establish the motion time series of the customer's head, gestures and limb movements. The system uses cameras to acquire the three-dimensional spatial coordinates of various joints on the customer's body, performs three-dimensional modeling of the joints, and uses Euclidean distance to calculate the spatial relationship between different joints to generate a skeleton structure model. Based on time series analysis, the system models the dynamic changes of the skeleton model to obtain the customer's standing posture, gait and posture change patterns. Temporal convolution processing is performed on the motion trajectory of joints established based on the 3D skeleton model, and the behavioral features of customers are extracted using a temporal convolutional network. Calculate the rate of change of customers' body posture in different time windows, and model the gait cycle, limb swing frequency, and head rotation angle to identify the trend of customers' business needs. Key point detection is performed on the customer's facial area to extract the coordinates of facial feature points, and feature changes under different expression states are calculated based on principal component analysis. Time series analysis is performed on the curvature of the corners of the mouth, the rise and fall of the eyebrows, and the opening and closing angle of the eyelids to calculate the rate of expression change and facial muscle movement characteristics, thereby obtaining the customer's emotional state.
4. The method for identifying customer behavior in electricity marketing based on multimodal analysis as described in claim 3, characterized in that: Temperature and humidity sensors, air quality detection modules, and noise monitoring equipment are deployed to acquire the temperature, humidity, carbon dioxide concentration, air flow rate, and environmental noise level of the business hall, and timestamps are added to the acquired data. Based on the collected temperature and humidity data, analyze the temperature and humidity fluctuations and extract the temperature and humidity change trends; The temperature and humidity data are compared with the customer's head movement trajectory, limb movement frequency and facial expression features extracted from the video data. The gait rate, dwell time, frequency of body micro-movements and changes in facial muscle activity of the customer under different temperature and humidity conditions are calculated. When the temperature and humidity exceed the set threshold and the customer's behavior pattern changes abnormally, the behavior confidence of the video data is adjusted to correct the misjudgment caused by temperature and humidity fluctuations and the customer behavior recognition results are corrected. Based on air quality data, the impact of carbon dioxide concentration and air flow rate on customer behavior patterns is analyzed; combined with the speech temporal characteristics of speech data extraction methods, the impact of air quality on the speech rate, volume, pitch stability and clarity of customer speech signals is analyzed. The degree of interference of air quality on customer voice feature parameters is calculated, and a compensation algorithm is used to correct the voice signal affected by the environment. Based on data from noise monitoring equipment, the noise interference coefficient in the business hall is calculated, and combined with the beamforming processing results of the voice data extraction method, the signal-to-noise ratio change of the voice signal is judged. When the ambient noise exceeds the set threshold, the beam direction of the microphone array is adjusted to enhance the signal strength of the target customer's voice, and the confidence level of the voice data is dynamically adjusted.
5. The method for identifying customer behavior in electricity marketing based on multimodal analysis as described in claim 4, characterized in that: The feature extraction method using cross-modal feature alignment includes assigning timestamps to the target customer's voice data and customer's video data respectively, and performing synchronization alignment based on a global clock so that the voice data and video data at the same moment have a unified time identifier. Based on synchronized voice and video data, a cross-modal Transformer feature alignment model is constructed to calculate the temporal correspondence between features of different modalities and extract highly correlated features. Based on the time synchronization information of voice data and video data, establish the correspondence between voice expression patterns and body movement patterns; Based on a feature-level fusion network, the emotional features of voice data and the behavioral features of video data are mapped in a high dimension to generate a fused customer behavior vector.
6. The method for identifying customer behavior in electricity marketing based on multimodal analysis as described in claim 5, characterized in that: The real-time analysis includes nonlinear feature mapping based on stacked autoencoders, and extracting features with representation capabilities higher than a set threshold by dimensionality reduction, and removing redundant information. A self-attention mechanism is used to calculate the feature weights of the behavior vector at different time steps and extract key behavior patterns. A bidirectional long short-term memory network is used to perform time series modeling on the fused customer behavior vector, calculate the temporal correlation of behavior patterns, and model historical behavior trends to enhance the ability to identify continuous behavior. A cross-modal fusion mechanism is constructed in the output layer of a bidirectional long short-term memory network. The cross-modal information entropy of behavioral patterns is calculated based on a multilayer perceptron, and the current behavioral state of customers is identified through a nonlinear classification method. Deep reinforcement learning is used to optimize model parameters, with the accuracy of customer behavior recognition as the reward signal, and the model hyperparameters are dynamically adjusted through a policy gradient optimization algorithm.
7. The method for identifying customer behavior in electricity marketing based on multimodal analysis as described in claim 6, characterized in that: The prediction of subsequent customer behavior includes using time series forecasting methods to construct a time series model of customer behavior and using gated loop units to predict the trend of customer behavior changes at different time steps. The system calculates the temporal variation trend of the fused customer behavior vector, uses convolutional temporal networks to analyze customer behavior patterns in different time windows, and predicts whether customers will continue to consult, leave the business hall, or escalate their emotions.
8. A system employing the multimodal analysis-based customer behavior identification method for electricity marketing as described in any one of claims 1 to 7, characterized in that, include: The data acquisition module is used to deploy data acquisition equipment to collect on-site data in real time and obtain multimodal data of the customer's voice, video and environment; The feature fusion module is used to perform time synchronization alignment on the collected multimodal data and to extract features through cross-modal feature alignment methods. as well as, The behavior recognition module is used to input the fused customer behavior vectors into the customer behavior recognition model for real-time analysis, to determine the customer's business needs, predict the customer's subsequent behavior, and optimize business processes based on customer behavior data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the electricity marketing customer behavior identification method based on multimodal analysis as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the electricity marketing customer behavior identification method based on multimodal analysis as described in any one of claims 1 to 7.