Interactive projection system based on image acquisition and rendering technology
By using image acquisition and rendering technology in the interactive projection system, users' image, voice and posture data are analyzed in real time and personalized interactive feedback is generated, which solves the problem that existing systems cannot generate personalized feedback based on users' emotional state, and achieves higher user experience and emotional analysis accuracy.
Patent Information
- Application Number
- CN202510149978.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Although the existing interactive projection system has interactive capabilities, it cannot generate personalized interactive feedback based on the user's emotional state, and its functionality needs to be improved.
Using an interactive projection system based on image acquisition and rendering technology, through the image acquisition module, expression recognition and emotion analysis module, speech recognition and emotion analysis module, posture recognition and instruction analysis module, instruction execution and feedback module, and rendering output module, the user's image, voice and posture data are captured and analyzed in real time to generate personalized interactive feedback.
It can generate highly personalized interactive feedback based on the user's emotional state, improve user experience, and improve the accuracy of sentiment analysis through multimodal fusion technology.
Smart Images

Figure CN120161940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction projection systems, and particularly to an interactive projection system based on image acquisition and rendering technologies. Background Art
[0002] Projection systems are widely used in life; these systems generally include a projector, a screen, and necessary connection cables and control devices. Users transmit content to the projector through a computer or other playback device, and then the projector projects images or videos onto the screen. The main advantages of such systems are relatively low cost, simple operation, and suitability for the display of static content.
[0003] However, with the development of technology and the improvement of user requirements, traditional projection systems have gradually revealed some limitations. For example, the interactivity of such systems is poor, and users cannot directly interact with the system.
[0004] After retrieval, the application solution with the Chinese patent application number CN202220672123.X discloses an interactive projection system, including a projector and a touch interactive light source device. The projector includes a projection optical machine and a camera module. The projection optical machine is used to emit a projection light beam towards the projection surface to form a projection range on the projection surface. The touch interactive light source device includes at least one infrared light source module and a wireless signal transmission module. The infrared light source module is used to generate an infrared light light curtain parallel to the projection surface and is electrically connected to the wireless signal transmission module. The projector is connected to the touch interactive light source device through the wireless signal transmission module, so that the touch interactive light source device is placed at the recommended installation position and the interactive projection system is in the interactive mode. The interactive projection system in the above patent has the following deficiencies: Although it has the ability to interact, it cannot generate personalized interactive feedback according to the emotional state of the user, and its functionality still needs to be improved. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose an interactive projection system based on image acquisition and rendering technologies.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] An interactive projection system based on image acquisition and rendering technologies, including:
[0008] An image acquisition module, which captures user images and video data and uses a high-definition camera and an infrared sensor for data acquisition;
[0009] An expression recognition and emotion analysis module, which uses deep learning algorithms to perform real-time analysis and recognition of the user's facial expressions and judge the user's emotional state;
[0010] The speech recognition and emotion analysis module uses natural language processing technology to transcribe and analyze the user's speech input in real time, and judge their emotional state and intention;
[0011] The gesture recognition and instruction parsing module uses a skeleton tracking algorithm to detect and recognize the user's body gestures in real time; according to the preset gesture templates, it parses the user's action intention and generates corresponding control instructions;
[0012] The instruction execution and feedback module converts the control instructions into system operations to drive the projection device to perform corresponding actions;
[0013] The rendering output module uses graphics rendering technology to dynamically generate projection content according to system instructions;
[0014] The data processing and storage module stores and manages all the collected data.
[0015] Preferably: The image acquisition module uses the OpenCV library for image capture and preprocessing, including denoising, enhancement, and cropping operations; it communicates with the camera device through a USB or network interface to achieve real-time data transmission;
[0016] The image acquisition module includes:
[0017] The camera sub-module, the camera sub-module is a multi-camera array, and the camera sub-module is equipped with a high-definition camera and an infrared sensor. High-definition camera: used to capture the user's face and body images, supporting high-resolution and high-frame-rate shooting;
[0018] The image preprocessing sub-module uses a filtering algorithm to remove image noise; it improves the image quality through contrast adjustment and edge enhancement.
[0019] Preferably: The expression recognition and emotion analysis module includes:
[0020] The facial feature extraction sub-module uses a pre-trained model such as Dlib or FaceNet to extract key facial feature points; it converts the key point data into feature vectors;
[0021] The expression classification sub-module constructs a deep learning model to classify facial expressions; it trains the model using labeled data and improves the generalization ability through data augmentation and transfer learning;
[0022] The emotion analysis sub-module combines facial expressions to comprehensively judge the user's emotional state;
[0023] The expression recognition and emotion analysis module uses a pre-trained model to extract key facial feature points, specifically as follows:
[0024] Select a pre-trained model: Select a pre-trained face keypoint detection model of Dlib or FaceNet, and accurately identify the key feature points of the face based on the model;
[0025] Load the model: Use the deep learning frameworks TensorFlow or PyTorch to load the pre-trained Dlib or FaceNet model;
[0026] Image preprocessing: Preprocess the input image, including scaling, cropping, and normalization, to ensure that the image meets the input requirements of the model;
[0027] Keypoint detection: Input the preprocessed image into the pre-trained model to obtain the keypoint coordinates of the face;
[0028] Feature vector generation: Convert the detected keypoint coordinates into feature vectors.
[0029] Preferably: The expression recognition and emotion analysis module trains the model using labeled data, specifically as follows:
[0030] Data collection and annotation: Collect facial expression images and perform manual annotation; the annotation content includes expression categories;
[0031] Data augmentation: Use data augmentation techniques, including random cropping, rotation, flipping, and color transformation, to increase the diversity of training data;
[0032] Model construction: Use the deep learning frameworks TensorFlow or PyTorch to construct a convolutional neural network model;
[0033] Model training: Input the labeled data into the convolutional neural network model for training; use the cross-entropy loss function and the Adam optimizer to adjust the model parameters through backpropagation.
[0034] Preferably: The expression recognition and emotion analysis module optimizes emotion prediction using time series analysis methods, specifically as follows:
[0035] Real-time emotion judgment: Use the trained facial expression classification model to judge the emotion of the real-time captured facial images;
[0036] Historical emotion data storage: Store the emotion judgment results of each time into the database to form a historical emotion data record of the user;
[0037] Time series analysis: Use ARIMA or LSTM to analyze the user's historical emotion data and identify the trends and patterns of emotion changes.
[0038] Preferably: The speech recognition and emotion analysis module includes:
[0039] Speech transcription sub-module, automatic speech recognition, using the SpeechRecognition library or Google Speech API, to convert speech into text;
[0040] Speech emotion analysis sub-module, by analyzing the changes in intonation and rhythm, to judge the user's emotional state, using NLP toolkits spaCy or NLTK, to parse the text content, extract emotion keywords, and comprehensively judge the user's emotion by combining speech and text information;
[0041] The speech recognition and emotion analysis module parses the text content to extract emotion keywords, and the specific method is as follows:
[0042] Speech transcription: Using automatic speech recognition technology, convert the user's speech into text;
[0043] Intonation and rhythm analysis: Analyze the changes in intonation and rhythm in the speech signal, and extract audio features;
[0044] NLP processing: Using NLP toolkits spaCy or NLTK, perform word segmentation, annotation, and syntactic analysis on the transcribed text; extract emotion keywords;
[0045] Multi-modal emotion judgment: Combine speech features and text content, and use an emotion analysis model to judge the user's emotional state.
[0046] Preferably: The posture recognition and instruction parsing module includes:
[0047] Human pose estimation sub-module, using the OpenPose or MediaPipe framework, to perform human pose estimation and skeleton tracking; calculate joint angles according to the skeleton point coordinates for action recognition;
[0048] Action classification sub-module, to match and classify continuous action sequences; use support vector machine or random forest algorithm to classify static postures;
[0049] Instruction generation sub-module, according to the preset posture templates and rules, parse the user's action intention, and generate corresponding control instructions; support users to customize posture templates to meet the needs of different application scenarios;
[0050] The posture recognition and instruction parsing module matches and classifies continuous action sequences, specifically as follows:
[0051] Action sequence capture: Use a camera to capture the user's continuous action sequence and input the video frames into the system;
[0052] Skeleton extraction: Use the human pose estimation algorithms OpenPose or MediaPipe to extract the human skeleton key points in each frame of the image;
[0053] Action feature extraction: Extract action features from the skeleton key points;
[0054] Action matching: Match and classify continuous action sequences; Based on DTW, compare action sequences of different lengths to find the most similar action patterns; Classify static postures based on SVM;
[0055] Instruction generation: Generate corresponding control instructions according to the matching results.
[0056] Preferably: The instruction execution and feedback module includes:
[0057] Instruction execution sub-module, which calls the operating system functions or APIs to drive the projection device to perform corresponding actions; Controls external devices through the GPIO interface;
[0058] Real-time feedback sub-module, which uses the audio synthesis libraries pygame.mixer or pydub to generate real-time sound effect feedback; Displays dynamic graphics or animations through the projection device.
[0059] Preferably: It also includes:
[0060] Multi-modal emotion fusion module: Adopts multi-modal fusion technology to weight and fuse the analysis results of facial expressions, speech emotions, and posture emotions; More accurately judge the overall emotional state of the user through the comprehensive emotional score after fusion;
[0061] The multi-modal emotion fusion module, the fusion method includes the following steps:
[0062] Data preparation: Collect a multi-modal emotion analysis data set containing facial expression, speech, and posture data;
[0063] Feature extraction: Extract features from each type of modal data respectively to obtain their respective feature vectors;
[0064] Feature fusion: Select a suitable fusion strategy according to the above method to fuse the feature vectors of different modalities;
[0065] Model training: Use the fused feature vectors to train an emotion classification model;
[0066] Model evaluation: Use cross-validation or other evaluation methods to evaluate the model and optimize the model parameters;
[0067] Real-time application: Deploy the trained model to a real-time system for emotion analysis of newly collected data.
[0068] Preferably: For the data-level fusion of the multimodal emotion fusion module, different modal data are used as inputs, sharing the parameter vector w, and the loss function L is calculated as follows:
[0069]
[0070] where y ij is the j-th modal label of the i-th sample, λ j is the weight coefficient, φ(x i ) is the feature vector of the i-th sample, d j is the parameter vector of the j-th modality, and Ω(w) is the regularization term.
[0071] The beneficial effects of the present invention are as follows:
[0072] 1. The system of the present invention can generate highly personalized interactive feedback according to the user's emotional state, improving the user experience; by adopting the multimodal fusion technology, the analysis results of facial expressions, speech emotions, and posture emotions are weighted and fused, improving the accuracy of emotion analysis.
[0073] 2. The present invention uses a deep learning framework to construct a convolutional neural network model for classifying and recognizing emotions in facial expressions; by using a pre-trained face key point detection model, facial feature points are extracted to improve the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 FIG. is a flowchart of the implementation of an interactive projection system based on image acquisition and rendering technology proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] The technical solutions of the present invention will be further described in detail below in conjunction with the specific embodiments.
[0076] Embodiment 1:
[0077] An interactive projection system based on image acquisition and rendering technology, comprising:
[0078] An image acquisition module that captures user images and video data, and uses a high-definition camera and an infrared sensor for data acquisition;
[0079] An expression recognition and emotion analysis module that, through deep learning algorithms, performs real-time analysis and recognition of the user's facial expressions, determines the user's emotional state, such as happy, angry, sad, etc.; in combination with historical data, further optimizes the accuracy of emotion analysis;
[0080] A speech recognition and emotion analysis module that uses natural language processing technology to transcribe and analyze the user's speech input in real time, and determines its emotional state and intention;
[0081] The posture recognition and instruction parsing module uses a skeleton tracking algorithm to detect and recognize the user's body posture in real time; according to the preset posture templates, it parses the user's action intentions and generates corresponding control instructions;
[0082] The instruction execution and feedback module converts the control instructions into system operations and drives the projection device to perform corresponding actions;
[0083] The rendering output module uses graphics rendering technology to dynamically generate projection content according to system instructions;
[0084] The data processing and storage module stores and manages all the collected data.
[0085] Among them, the image acquisition module uses the OpenCV library for image capture and preprocessing, including operations such as denoising, enhancement, and cropping; it communicates with the camera device through a USB or network interface to achieve real-time data transmission;
[0086] The image acquisition module includes:
[0087] The camera sub-module, which is a multi-camera array. The camera sub-module is equipped with a high-definition camera and an infrared sensor. The high-definition camera is used to capture the user's facial and body images and supports high-resolution and high-frame-rate shooting;
[0088] The image preprocessing sub-module uses a filtering algorithm to remove image noise; it improves the image quality through contrast adjustment and edge enhancement.
[0089] Among them, the facial expression recognition and emotion analysis module constructs a convolutional neural network model using a deep learning framework (such as TensorFlow or PyTorch) to classify facial expressions and recognize emotions; it uses a pre-trained face key point detection model (such as Dlib or FaceNet) to extract facial feature points to improve the recognition accuracy; according to the recognition results, it calls an emotion analysis algorithm (such as VADER or BERT) to judge the user's emotional state;
[0090] The facial expression recognition and emotion analysis module includes:
[0091] The facial feature extraction sub-module uses a pre-trained model (such as Dlib or FaceNet) to extract facial key feature points; it converts the key point data into feature vectors for subsequent classification;
[0092] The facial expression classification sub-module constructs a deep learning model to classify facial expressions (such as happy, angry, sad, etc.); it trains the model using labeled data and improves the generalization ability through data augmentation and transfer learning;
[0093] The emotion analysis sub-module combines facial expressions to comprehensively judge the user's emotional state; uses time series analysis methods and combines the user's historical emotion data to further optimize emotion prediction.
[0094] Among them, the expression recognition and emotion analysis module uses a pre-trained model to extract key facial feature points, specifically as follows:
[0095] Select a pre-trained model: Select a pre-trained face key point detection model such as Dlib or FaceNet; accurately identify the key feature points of the face (such as eyes, nose, mouth, chin, etc.) based on the model;
[0096] Load the model: Use a deep learning framework (such as TensorFlow or PyTorch) to load the pre-trained Dlib or FaceNet model;
[0097] Image preprocessing: Preprocess the input image, including scaling, cropping, and normalization, to ensure that the image meets the input requirements of the model;
[0098] Key point detection: Input the preprocessed image into the pre-trained model to obtain the coordinates of the key points of the face; for example, Dlib will return the coordinates of 68 key points;
[0099] Feature vector generation: Convert the detected key point coordinates into feature vectors; this can be achieved by calculating the relative positions, angles, or other geometric features between the key points; the feature vectors are used for subsequent expression classification and emotion analysis.
[0100] Among them, the expression recognition and emotion analysis module uses labeled data to train the model, specifically as follows:
[0101] Data collection and annotation: Collect a large number of facial expression images and perform manual annotation; the annotation content includes expression categories (such as happy, angry, sad, etc.);
[0102] Data augmentation: To improve the generalization ability of the model, use data augmentation techniques such as random cropping, rotation, flipping, and color transformation to increase the diversity of the training data;
[0103] Model construction: Use a deep learning framework (such as TensorFlow or PyTorch) to construct a convolutional neural network (CNN) model;
[0104] Model training: Input the labeled data into the CNN model for training; use the cross-entropy loss function and the Adam optimizer to adjust the model parameters through backpropagation;
[0105] Transfer learning: If the training data is insufficient, transfer learning techniques can be used; first, a base network is pre-trained on a large dataset (such as ImageNet), and then fine-tuned on the facial expression dataset.
[0106] Among them, the expression recognition and emotion analysis module optimizes emotion prediction using time series analysis methods as follows:
[0107] Real-time emotion judgment: Use the trained facial expression classification model to judge the emotion of the face images captured in real time;
[0108] Historical emotion data storage: Store the results of each emotion judgment in the database to form a historical emotion data record of the user;
[0109] Time series analysis: Use time series analysis methods (such as ARIMA, LSTM, etc.) to analyze the user's historical emotion data and identify the trends and patterns of emotion changes;
[0110] Emotion prediction optimization: Combine the current emotion judgment results and historical emotion data, and use machine learning algorithms (such as linear regression, random forest, etc.) to further optimize the accuracy of emotion prediction.
[0111] Among them, the speech recognition and emotion analysis module includes:
[0112] Speech transcription sub-module, automatic speech recognition (ASR), using the SpeechRecognition library or Google Speech API, to convert speech into text;
[0113] Speech emotion analysis sub-module, by analyzing the changes in tone and rhythm, judge the user's emotional state, use NLP toolkits (such as spaCy or NLTK) to parse the text content, extract emotion keywords, and combine speech and text information to comprehensively judge the user's emotion.
[0114] Among them, the speech recognition and emotion analysis module extracts emotion keywords by parsing the text content in the following specific ways:
[0115] Speech transcription: Use automatic speech recognition (ASR) technology to convert the user's speech into text; common tools include the SpeechRecognition library or Google Speech API;
[0116] Tone and rhythm analysis: Analyze the changes in tone and rhythm in the speech signal and extract audio features (such as frequency, energy, duration, etc.); these features can reflect the user's emotional state;
[0117] NLP Processing: Use NLP toolkits (such as spaCy or NLTK) to tokenize, annotate, and perform syntactic analysis on the transcribed text; extract sentiment keywords such as "happy", "angry", "sad", etc.
[0118] Multimodal Sentiment Judgment: Combine speech features and text content, and use a sentiment analysis model to judge the user's sentiment state.
[0119] Among them, the posture recognition and instruction parsing module includes:
[0120] Human Pose Estimation Sub-module: Use the OpenPose or MediaPipe framework to perform human pose estimation and skeleton tracking; calculate joint angles based on skeleton point coordinates for action recognition.
[0121] Action Classification Sub-module: Match and classify continuous action sequences; use algorithms such as Support Vector Machine (SVM) or Random Forest to classify static postures.
[0122] Instruction Generation Sub-module: Parse the user's action intention according to preset posture templates and rules, and generate corresponding control instructions; support users to customize posture templates to meet the needs of different application scenarios.
[0123] Among them, the posture recognition and instruction parsing module matches and classifies continuous action sequences as follows:
[0124] Action Sequence Capture: Use a camera to capture the user's continuous action sequence and input the video frames into the system.
[0125] Skeleton Extraction: Use a human pose estimation algorithm (such as OpenPose or MediaPipe) to extract the human skeleton key points in each frame of the image.
[0126] Action Feature Extraction: Extract action features from the skeleton key points, such as joint angles, speed, acceleration, etc.
[0127] Action Matching: Match and classify continuous action sequences; common methods include Dynamic Time Warping (DTW) and Support Vector Machine (SVM); DTW can compare action sequences of different lengths to find the most similar action pattern; SVM can classify static postures.
[0128] Instruction Generation: Generate corresponding control instructions according to the matching results; for example, when it is detected that the user makes a "waving" action, the system can generate an instruction to "switch pages".
[0129] Among them, the instruction execution and feedback module includes:
[0130] The instruction execution sub-module calls operating system functions or APIs to drive the projection device to perform corresponding actions; controls external devices (such as lights, speakers, etc.) through the GPIO interface;
[0131] The real-time feedback sub-module uses an audio synthesis library (such as pygame.mixer or pydub) to generate real-time sound effect feedback; displays dynamic graphics or animations through the projection device.
[0132] The working process of the interactive projection system includes the following steps:
[0133] S1: Initialization: After the system starts, each module loads the configuration file in sequence and performs self-checking;
[0134] S2: Data acquisition: The image acquisition module and the speech recognition module start to work, and capture user image and speech data in real time;
[0135] S3: Data processing: Based on facial expression recognition, speech recognition, and gesture recognition, the collected data is processed respectively to extract key features;
[0136] S4: Sentiment analysis: According to the processing results, the sentiment analysis module judges the user's emotional state and intention;
[0137] S5: Instruction generation: According to the sentiment analysis results, the instruction parsing module generates corresponding control instructions;
[0138] S6: Instruction execution: After receiving the instruction, the instruction execution module calls the corresponding function or API to drive the projection device to perform corresponding actions;
[0139] S7: Feedback output: Provide an instant interactive experience to the user through the projection device and the audio system;
[0140] S8: Data storage: Store all the collected data in the database.
[0141] Embodiment 2:
[0142] An interactive projection system based on image acquisition and rendering technology. On the basis of Embodiment 1, this embodiment further includes:
[0143] The multi-modal sentiment fusion module: Adopts multi-modal fusion technology to perform weighted fusion on the analysis results of facial expressions, speech emotions, and gesture emotions; more accurately judges the overall emotional state of the user through the comprehensive emotion score after fusion;
[0144] For the multi-modal sentiment fusion module, the fusion method includes the following steps:
[0145] Data preparation: Collect a multi-modal sentiment analysis data set containing facial expression, speech, and gesture data;
[0146] Feature extraction: Feature extraction is performed on each type of modal data respectively to obtain their respective feature vectors;
[0147] Feature fusion: According to the above method, a suitable fusion strategy is selected to fuse the feature vectors of different modalities;
[0148] Model training: Use the fused feature vectors to train a sentiment classification model (such as a support vector machine, random forest, or deep learning model);
[0149] Model evaluation: Use cross-validation or other evaluation methods to evaluate the model and optimize the model parameters;
[0150] Real-time application: Deploy the trained model to a real-time system to perform sentiment analysis on newly collected data;
[0151] For the data-level fusion of the multi-modal sentiment fusion module, different modal data are used as inputs, sharing the parameter vector w, and calculating the loss function L as follows:
[0152]
[0153] where y ij is the j-th modal label of the i-th sample, λ j is the weight coefficient, φ(x i ) is the feature vector of the i-th sample, d j is the parameter vector of the j-th modality, and Ω(w) is the regularization term.
[0154] As mentioned above, only the preferred specific embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An interactive projection system based on image acquisition and rendering technology, characterized in that: include: Image acquisition module, which captures user images and video data, using high-definition cameras and infrared sensors for data collection; The expression recognition and emotion analysis module uses deep learning algorithms to analyze and recognize users' facial expressions in real time and determine their emotional state; The speech recognition and sentiment analysis module uses natural language processing technology to transcribe and analyze the user's speech input in real time to determine their emotional state and intention; The posture recognition and command parsing module uses a skeleton tracking algorithm to detect and recognize the user's body posture in real time; according to the preset posture template, it analyzes the user's action intention and generates corresponding control instructions; The command execution and feedback module converts the control command into system operation and drives the projection device to perform corresponding actions; The rendering output module uses graphics rendering technology to dynamically generate projection content according to system instructions; The data processing and storage module stores and manages all collected data.
2. The interactive projection system based on image acquisition and rendering technology according to claim 1, characterized in that: The image acquisition module uses the OpenCV library to perform image capture and preprocessing, including denoising, enhancement and cropping operations; Communicate with camera devices via USB or network interface to achieve real-time data transmission; The image acquisition module comprises: Camera submodule: The camera submodule is a multi-camera array. The camera submodule is equipped with a high-definition camera and an infrared sensor. High-definition camera: used to capture user facial and body images, supporting high-resolution and high-frame rate shooting; Image preprocessing submodule, the image preprocessing submodule uses filtering algorithms to remove image noise; and improves image quality through contrast adjustment and edge enhancement.
3. The interactive projection system based on image acquisition and rendering technology according to claim 1, characterized in that: The expression recognition and emotion analysis module includes: The facial feature extraction submodule uses the pre-trained model Dlib or FaceNet to extract key facial feature points and convert key point data into feature vectors; The expression classification submodule builds a deep learning model to classify facial expressions; the model is trained using labeled data and generalization capabilities are improved through data augmentation and transfer learning; The emotion analysis submodule combines facial expressions to comprehensively judge the user's emotional state; The expression recognition and emotion analysis module uses a pre-trained model to extract key facial feature points, as follows: Select a pre-trained model: Select a Dlib or FaceNet pre-trained facial key point detection model to accurately identify the key feature points of the face based on the model; Loading model: Use deep learning frameworks such as TensorFlow or PyTorch to load pre-trained Dlib or FaceNet models; Image preprocessing: Preprocess the input image, including scaling, cropping, and normalization, to ensure that the image meets the input requirements of the model; Key point detection: Input the preprocessed image into the pre-trained model to obtain the key point coordinates of the face; Feature vector generation: Convert the detected key point coordinates into feature vectors.
4. The interactive projection system based on image acquisition and rendering technology according to claim 3, characterized in that: The expression recognition and emotion analysis module uses the labeled data to train the model, as follows: Data collection and annotation: Collect facial expression images and manually annotate them; the annotation content includes expression categories; Data augmentation: Use data augmentation techniques, including random cropping, rotation, flipping, and color transformation, to increase the diversity of training data; Model building: Use deep learning frameworks such as TensorFlow or PyTorch to build convolutional neural network models; Model training: Input the labeled data into the convolutional neural network model for training; use the cross entropy loss function and Adam optimizer to adjust the model parameters through back propagation.
5. The interactive projection system based on image acquisition and rendering technology according to claim 3, characterized in that: The expression recognition and emotion analysis module uses a time series analysis method to optimize emotion prediction, as follows: Real-time emotion judgment: Use the trained facial expression classification model to judge the emotion of facial images captured in real time; Historical emotion data storage: Each emotion judgment result is stored in the database to form the user's historical emotion data record; Time series analysis: Use ARIMA or LSTM to analyze users’ historical sentiment data and identify trends and patterns in sentiment changes.
6. The interactive projection system based on image acquisition and rendering technology according to claim 1, characterized in that: The speech recognition and sentiment analysis module includes: Speech transcription submodule, automatic speech recognition, using the SpeechRecognition library or Google Speech API to convert speech into text; The speech emotion analysis submodule determines the user's emotional state by analyzing the changes in intonation and rhythm. It uses the NLP toolkit spaCy or NLTK to parse the text content, extract emotional keywords, and combine speech and text information to comprehensively determine the user's emotions. The speech recognition and sentiment analysis module parses the text content to extract sentiment keywords in the following way: Voice transcription: Use automatic speech recognition technology to convert user speech into text; Intonation and rhythm analysis: Analyze the intonation and rhythm changes in speech signals and extract audio features; NLP processing: Use NLP toolkits such as spaCy or NLTK to segment, annotate, and perform syntactic analysis on the transcribed text; extract sentiment keywords; Multimodal sentiment judgment: Combine speech features and text content and use sentiment analysis models to judge the user's emotional state.
7. The interactive projection system based on image acquisition and rendering technology according to claim 1, characterized in that: The gesture recognition and instruction parsing module includes: The human pose estimation submodule uses the OpenPose or MediaPipe framework to perform human pose estimation and skeleton tracking. It calculates joint angles based on the coordinates of skeleton points for action recognition. The action classification submodule matches and classifies continuous action sequences; it uses support vector machines or random forest algorithms to classify static postures; The instruction generation submodule analyzes the user's action intention and generates corresponding control instructions based on the preset posture templates and rules; it supports users to customize posture templates to meet the needs of different application scenarios; The gesture recognition and instruction parsing module matches and classifies continuous action sequences, as follows: Motion sequence capture: Use a camera to capture the user's continuous motion sequence and input the video frames into the system; Skeleton extraction: Use the human pose estimation algorithm OpenPose or MediaPipe to extract the key points of the human skeleton in each frame of the image; Action feature extraction: extract action features from skeleton key points; Action matching: Match and classify continuous action sequences; compare action sequences of different lengths based on DTW to find the most similar action pattern; classify static postures based on SVM; Instruction generation: Generate corresponding control instructions based on the matching results.
8. The interactive projection system based on image acquisition and rendering technology according to claim 1, characterized in that: The instruction execution and feedback module includes: The instruction execution submodule calls the operating system function or API to drive the projection device to perform corresponding actions; it controls external devices through the GPIO interface; The real-time feedback submodule uses the audio synthesis library pygame.mixer or pydub to generate real-time sound feedback; dynamic graphics or animations are displayed through a projection device.
9. An interactive projection system based on image acquisition and rendering technology according to any one of claims 1 to 8, characterized in that: Also includes: Multimodal emotion fusion module: adopts multimodal fusion technology to weightedly fuse the analysis results of facial expression, voice emotion and posture emotion; through the fused comprehensive emotion score, the overall emotional state of the user can be judged more accurately; The multimodal emotion fusion module and the fusion method include the following steps: Data preparation: Collect a multimodal sentiment analysis dataset containing facial expression, speech, and gesture data; Feature extraction: Extract features from each modal data to obtain their respective feature vectors; Feature fusion: Select a suitable fusion strategy based on the above method to fuse feature vectors of different modalities; Model training: Use the fused feature vectors to train the sentiment classification model; Model evaluation: Use cross-validation or other evaluation methods to evaluate the model and optimize model parameters; Real-time application: Deploy the trained model to a real-time system to perform sentiment analysis on newly collected data.
10. The interactive projection system based on image acquisition and rendering technology according to claim 9, characterized in that: The data-level fusion of the multimodal emotion fusion module takes data of different modalities as input, shares the parameter vector w, and calculates the loss function L, as follows: Among them, y ij is the jth modality label of the ith sample, λ j is the weight coefficient, φ(x i ) is the feature vector of the i-th sample, d j is the parameter vector of the jth mode, and Ω(w) is the regularization term.
Citation Information
Patent Citations
Interactive projection system
CN217060959U
Cited By
Emotion recognition method and emotion analysis system
CN120472435A