Video processing method based on AI support
Through AI-based video processing methods, the problems of low image quality repair efficiency, insufficient content understanding accuracy, cumbersome editing process and low security detection efficiency in traditional video processing technologies are solved, and efficient video quality repair, content understanding, editing and security detection are achieved, improving user experience.
Patent Information
- Application Number
- CN202510666331.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-15
AI Technical Summary
When traditional video processing technology faces complex video processing needs, there are problems such as low image quality repair efficiency, insufficient content understanding accuracy, cumbersome editing process and low content security detection efficiency.
AI-based video processing methods are adopted, including data acquisition, preprocessing, AI processing, business logic execution and user interaction, deep learning algorithms and generative adversarial networks are used for video repair and enhancement, combined with computer vision and natural language processing for content understanding and analysis, and introduced intelligent material screening and multimodal fusion model for content security detection.
It significantly improves the efficiency and accuracy of video quality repair, improves the accuracy of content understanding and editing efficiency, realizes efficient content security detection, reduces labor costs and improves user interaction experience.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and specifically to a video processing method based on AI support. Background Art
[0002] With the rapid development of digital information, video, as an important carrier of information dissemination, is seeing its application scenarios continue to expand, covering a wide range of fields including entertainment, education, security, and healthcare. However, traditional video processing technologies are gradually revealing many limitations when faced with the growing demand for complex video processing, as shown below:
[0003] 1. Video quality restoration and enhancement: Many old videos suffer from blurry images, excessive noise, and color distortion due to factors such as filming equipment and storage time. For example, the original film stock of some classic films has become scratched and faded during long-term storage, resulting in a poor visual experience for viewers. Traditional restoration methods often rely on manual frame-by-frame processing, which is labor-intensive and has limited effectiveness.
[0004] 2. Video content understanding and analysis: In security surveillance scenarios, it is necessary to quickly and accurately identify abnormal behavior, such as intrusions and fights, from large amounts of real-time surveillance video. Traditional video analysis technologies are mostly based on simple rule matching, which makes it difficult to accurately identify abnormal behavior in complex scenarios.
[0005] 3. Improved video editing and creation efficiency: Traditional video editing workflows are cumbersome, requiring professional editors to devote significant time and effort, from screening and editing footage to adding special effects and final synthesis. For example, creating a promotional short video might require an editor to manually select the appropriate clips from a vast library of footage, then individually edit and splice them, add subtitles and music, and so on. This entire process can take hours or even days.
[0006] 4. Video Content Safety Detection: With the explosive growth of online video content, inappropriate content (such as violence, pornography, and false information) has also increased. For example, on some social media platforms, some users may post videos containing content that violates laws, regulations, or ethical standards. Traditional content review methods, which primarily rely on manual review, are not only inefficient and unable to cope with the massive amount of video content, but are also prone to inconsistent review standards.
[0007] Based on the above, a video processing method based on AI support is invented. Summary of the Invention
[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0009] A video processing method based on AI support includes the following specific steps:
[0010] S1, data acquisition: obtaining video data from various data sources;
[0011] S2, data preprocessing: preliminary processing of the acquired raw video data to meet the requirements of subsequent AI processing. The preliminary processing includes video format conversion, denoising, frame extraction, and compression;
[0012] S3, AI processing: First, based on the specific video processing task, the corresponding AI model is called for processing, and then the output of the AI model is integrated. The AI model includes object detection and recognition, video restoration and enhancement, video content understanding and analysis, video editing and creation, and video content security detection;
[0013] S4, business logic: Based on business needs and AI-processed data, specific business logic operations are performed;
[0014] S5, user interaction: providing an operation interface for users;
[0015] S6, intelligent interaction: Based on user interaction, it can interact with users through voice, gestures, and eye contact to optimize the user's interactive experience.
[0016] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of target detection and recognition are: first calling the target detection algorithm model, then inputting the video frame data into the model, so that the model can extract features of the image through a convolutional neural network, and at the same time, using the region generation network to generate candidate regions that may contain targets, and then through classification and regression operations, identify the targets in the video frames and determine their categories and location information.
[0017] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of video restoration and enhancement are: first calling the generative adversarial network model, then using the generator to receive the damaged video frame image as input, and by learning the features of a large number of high-quality images, trying to generate a repaired image, and then the discriminator discriminates between the generated image and the real high-quality image, the two are trained against each other, and the parameters of the generator are continuously iterated to make the generated repaired image close to the real high-quality image, thereby realizing the restoration and enhancement of video quality.
[0018] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of the video content understanding and analysis are: first, using computer vision technology such as image segmentation and target tracking to identify and analyze the characters, objects, and scene elements in the video, and at the same time, performing speech recognition and converting the audio in the video into text, and then using natural language processing technology to perform sentiment analysis and topic classification on the text, and then fusing the visual and language information to understand the semantic information in the video, thereby realizing the classification, sentiment analysis, and behavior analysis operations of the video content.
[0019] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of video editing and creation are: first calling the intelligent algorithm model, then automatically screening relevant materials according to the video theme and style requirements set by the user, and then editing and splicing the screened materials through time series analysis and editing algorithms, and adding corresponding special effects, subtitles and music.
[0020] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of the video content security detection are: first calling the multimodal fusion deep learning model, and then analyzing the video image, audio, and text information to analyze whether there is any bad content in the video.
[0021] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of S4 are:
[0022] S41, receiving and processing: receiving data from AI processing and integration;
[0023] S42, Business Requirements Analysis: Based on the business requests submitted by users during user interaction or the business rules preset by the system, the business requirements are analyzed in depth to clarify the specific goals, requirements and constraints of the current business tasks;
[0024] S43, Data Processing and Integration: Further data processing and integration of received AI-processed data based on business needs;
[0025] S44, business logic execution: Execute specific business logic operations based on the analyzed business requirements and the processed and integrated data. The business logic operations include security monitoring services, video editing services, video content review services, and business data analysis services.
[0026] S45, result packaging and output: First, the results after the business logic execution are packaged and converted into a format that can be recognized and displayed by the user interaction. Then, the packaged results are output to the user interaction for the user to view, download or further operate.
[0027] As a preferred embodiment of the AI-supported video processing method described in the present invention, the specific steps of the security monitoring service are as follows: when the AI processing detects that there are situations in the video that meet the abnormal behavior alarm rules, the business logic triggers the alarm mechanism, sends an alarm message to relevant personnel, and stores the video clip containing the abnormal behavior in a designated storage location for subsequent review and tracing;
[0028] The specific steps of the video editing business are: synthesizing the integrated video materials, special effects, subtitles and music content according to the set logic and sequence to generate the final video work;
[0029] The specific steps of the video content review service are: review and process suspected inappropriate content detected by AI processing according to the criteria for determining inappropriate content; for videos determined to be inappropriate content, remove them from shelves, block them, and delete them according to the prescribed process, and record relevant processing logs; for content that cannot be determined, submit it for manual review;
[0030] The specific steps of the commercial data analysis business are: conducting commercial data analysis on the analysis results of video data based on AI processing, calculating key indicators, generating data analysis reports, and providing data support for merchants' decision-making.
[0031] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of S5 are:
[0032] S51, user input: the user submits an operation request through various terminals in user interaction;
[0033] S52, request parsing: After receiving user input, the request can be parsed to understand the user's specific needs;
[0034] S53, interface display preparation: while waiting for the business logic to process the request, the interface can be prepared for display and the interface display content can be updated according to the request type;
[0035] S54, result reception and analysis: After the business logic completes processing and outputs the result to the user interface, the user interface can first receive the processing result, then parse the result to determine the type and content of the result, and then prepare the corresponding display method according to the different result types;
[0036] S55, results display:
[0037] S551, video result display: If the processing result is a video file, a video playback window is provided on the interface, supporting various playback control functions. At the same time, a download button is provided to facilitate users to save the video locally, and a sharing function is provided to enable users to share the video to social media platforms with one click;
[0038] S552, data result display: The data results of the data analysis report are displayed on the interface in the form of intuitive charts and tables, so that users can quickly understand the meaning of the data.
[0039] As a preferred solution of the AI-supported video processing method described in the present invention, the specific steps of S6 are:
[0040] S61, User Input Capture: Utilizes multiple sensors and input devices to capture user input information in real time;
[0041] S62, Input Information Preprocessing: Preprocess the captured raw input information to improve the accuracy and efficiency of subsequent processing; perform noise reduction on the voice data to remove environmental noise interference and enhance the clarity of the voice signal; use endpoint detection technology to determine the start and end positions of the voice and segment valid voice segments; perform image enhancement and feature extraction on gesture and eye contact interaction data;
[0042] S63, Command Parsing and Intent Recognition: Utilizes natural language processing and computer vision technologies to deeply analyze pre-processed input information. For voice command processing, after converting speech to text, the system uses a semantic analysis model to understand the user's intended meaning and identify the user's action. For gesture and eye contact operations, the system determines the user's intended action based on pre-defined gesture rules and eye contact interaction logic.
[0043] S64, Intelligent Response Generation: Generates personalized intelligent responses based on identified user intent, combined with historical user behavior data and preference models. At the same time, it incorporates affective computing technology into the response process, adjusting the response method and content based on the user's current emotional state, providing more emotionally sensitive feedback.
[0044] S65, interactive feedback execution: Feedback the generated intelligent response to the user in an appropriate manner to achieve human-computer interaction.
[0045] Compared with existing technologies:
[0046] 1. Targeting video quality restoration and enhancement: This invention leverages deep learning algorithms and generative adversarial networks to automatically learn the characteristics of large amounts of high-quality video data. For example, in the restoration of classic films, AI can intelligently identify flaws such as scratches and discoloration on film. Using algorithms, it repairs the damaged areas, significantly improving video clarity and color reproduction. Compared to traditional methods, this method is tens or even hundreds of times more efficient, with more stable restoration quality. This allows old videos to be revitalized, providing users with a better visual experience.
[0047] 2. Regarding the problem of video content understanding and analysis: The present invention is based on AI-based video content analysis technology, which integrates computer vision and deep learning. It can perform in-depth semantic understanding of elements such as people, objects, and actions in the video. In security monitoring, it can quickly and accurately identify abnormal behaviors such as people breaking into and fighting, and trigger alarms in time. Compared with manual screening, the efficiency is significantly improved and the risk of omissions is reduced. In the retail industry, through the analysis of store surveillance videos, customer behavior habit data such as the area where customers stay and the order in which they browse products can be accurately obtained, providing a scientific basis for merchants to optimize store layout and product display, and tapping into commercial value that is difficult to discover with traditional technologies.
[0048] 3. Targeting the improvement of video editing and creation efficiency: This AI-based video processing solution introduces intelligent material screening, automatically matching relevant materials based on the video theme. The intelligent editing algorithm can quickly generate a variety of editing options and automatically add appropriate special effects, subtitles, and music. For example, producing a short promotional video used to take hours or even days. Now, with the help of AI technology, the initial creation can be completed in minutes, greatly shortening the video editing and creation cycle and reducing labor costs. It also provides creators with more creative inspiration and improves the efficiency and quality of content production.
[0049] 4. Regarding video content security detection: This AI-based video content security detection technology utilizes a multimodal deep learning model to comprehensively analyze video images, audio, text, and other information, enabling rapid and accurate identification of objectionable content such as violence, pornography, and false information. On social media platforms and other platforms, uploaded videos can be inspected in real time, promptly filtering out and addressing objectionable content. This improves efficiency by over a thousand times compared to manual review, effectively maintaining a healthy online video environment, protecting the legitimate rights and interests of users, and reducing platform operational risks. DETAILED DESCRIPTION
[0050] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below.
[0051] The present invention provides a video processing method based on AI support, comprising the following specific steps:
[0052] S1, data acquisition: obtaining video data from various data sources;
[0053] S2, data preprocessing: preliminary processing of the acquired raw video data to meet the requirements of subsequent AI processing. The preliminary processing includes video format conversion, denoising, frame extraction, and compression;
[0054] S3, AI processing: First, based on the specific video processing task, the corresponding AI model is called for processing, and then the output of the AI model is integrated. The AI model includes object detection and recognition, video restoration and enhancement, video content understanding and analysis, video editing and creation, and video content security detection;
[0055] The specific steps of target detection and recognition are as follows: first, calling the target detection algorithm model, then inputting the video frame data into the model so that the model can extract features from the image through a convolutional neural network, and at the same time, using a region generation network to generate candidate regions that may contain targets, and then through classification and regression operations, identifying the targets in the video frame and determining their categories and location information;
[0056] The specific steps of the video restoration and enhancement are as follows: first, a generative adversarial network model is called, and then a generator is used to receive damaged video frame images as input. By learning the features of a large number of high-quality images, the generator attempts to generate a restored image. Then, a discriminator discriminates between the generated image and the real high-quality image. The two are trained against each other, and the parameters of the generator are continuously iteratively optimized to make the generated restored image close to the real high-quality image, thereby achieving video quality restoration and enhancement;
[0057] The specific steps of the video content understanding and analysis are as follows: first, using computer vision technologies such as image segmentation and target tracking to identify and analyze characters, objects, and scene elements in the video; at the same time, performing speech recognition on the audio in the video and converting it into text; then using natural language processing technology to perform sentiment analysis and topic classification on the text; and then fusing visual and language information to understand the semantic information in the video, thereby achieving video content classification, sentiment analysis, and behavior analysis operations;
[0058] The specific steps of video editing and creation are as follows: first, the intelligent algorithm model is called, then the relevant materials are automatically screened according to the video theme and style requirements set by the user, and then the screened materials are edited and spliced through time series analysis and editing algorithms, and corresponding special effects, subtitles and music are added;
[0059] The specific steps of the video content security detection are: first call the multimodal fusion deep learning model, then analyze the video image, audio, and text information to analyze whether there is any bad content in the video
[0060] S4, business logic: Based on business needs and AI-processed data, specific business logic operations are performed;
[0061] The specific steps of S4 are:
[0062] S41, receiving and processing: receiving data from AI processing and integration;
[0063] S42, Business Requirements Analysis: Based on the business requests submitted by users during user interaction or the business rules preset by the system, the business requirements are analyzed in depth to clarify the specific goals, requirements and constraints of the current business tasks;
[0064] S43, Data Processing and Integration: Further data processing and integration of received AI-processed data based on business needs;
[0065] S44, business logic execution: Execute specific business logic operations based on the analyzed business requirements and the processed and integrated data. The business logic operations include security monitoring services, video editing services, video content review services, and business data analysis services.
[0066] The specific steps of the security monitoring business are as follows: when the AI processing detects that there are abnormal behavior alarm rules in the video, the business logic triggers the alarm mechanism, sends an alarm message to the relevant personnel, and stores the video clip containing the abnormal behavior in a designated storage location for subsequent review and tracing;
[0067] The specific steps of the video editing business are: synthesizing the integrated video materials, special effects, subtitles and music content according to the set logic and sequence to generate the final video work;
[0068] The specific steps of the video content review service are: review and process suspected inappropriate content detected by AI processing according to the criteria for determining inappropriate content; for videos determined to be inappropriate content, remove them from shelves, block them, and delete them according to the prescribed process, and record relevant processing logs; for content that cannot be determined, submit it for manual review;
[0069] The specific steps of the business data analysis business are: conducting business data analysis on the analysis results of video data based on AI processing, calculating key indicators, generating data analysis reports, and providing data support for merchants' decision-making;
[0070] S45, result packaging and output: First, the results after the business logic execution are packaged and converted into a format that can be recognized and displayed by the user interface. Then, the packaged results are output to the user interface for the user to view, download or further operate;
[0071] S5, user interaction: providing an operation interface for users;
[0072] The specific steps of S5 are:
[0073] S51, user input: the user submits an operation request through various terminals in user interaction;
[0074] S52, request parsing: After receiving user input, the request can be parsed to understand the user's specific needs;
[0075] S53, interface display preparation: while waiting for the business logic to process the request, the interface can be prepared for display and the interface display content can be updated according to the request type;
[0076] S54, result reception and analysis: After the business logic completes processing and outputs the result to the user interface, the user interface can first receive the processing result, then parse the result to determine the type and content of the result, and then prepare the corresponding display method according to the different result types;
[0077] S55, results display:
[0078] S551, video result display: If the processing result is a video file, a video playback window is provided on the interface, supporting various playback control functions. At the same time, a download button is provided to facilitate users to save the video locally, and a sharing function is provided to enable users to share the video to social media platforms with one click;
[0079] S552, data result display: The data results of the data analysis report are displayed on the interface in the form of intuitive charts and tables, which makes it easy for users to quickly understand the meaning of the data;
[0080] S6, intelligent interaction: Based on user interaction, it can interact with users through voice, gestures, and eye contact to optimize the user's interactive experience;
[0081] The specific steps of S6 are:
[0082] S61, User Input Capture: Utilizes multiple sensors and input devices to capture user input information in real time;
[0083] S62, Input Information Preprocessing: Preprocess the captured raw input information to improve the accuracy and efficiency of subsequent processing; perform noise reduction on the voice data to remove environmental noise interference and enhance the clarity of the voice signal; use endpoint detection technology to determine the start and end positions of the voice and segment valid voice segments; perform image enhancement and feature extraction on gesture and eye contact interaction data;
[0084] S63, Command Parsing and Intent Recognition: Utilizes natural language processing and computer vision technologies to deeply analyze pre-processed input information. For voice command processing, after converting speech to text, the system uses a semantic analysis model to understand the user's intended meaning and identify the user's action. For gesture and eye contact operations, the system determines the user's intended action based on pre-defined gesture rules and eye contact interaction logic.
[0085] S64, Intelligent Response Generation: Generates personalized intelligent responses based on identified user intent, combined with historical user behavior data and preference models. At the same time, it incorporates affective computing technology into the response process, adjusting the response method and content based on the user's current emotional state, providing more emotionally sensitive feedback.
[0086] S65, interactive feedback execution: Feedback the generated intelligent response to the user in an appropriate manner to achieve human-computer interaction.
[0087] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A video processing method based on AI support, characterized in that: The specific steps are as follows: S1, data acquisition: obtaining video data from various data sources; S2, data preprocessing: preliminary processing of the acquired raw video data to meet the requirements of subsequent AI processing. The preliminary processing includes video format conversion, denoising, frame extraction, and compression; S3, AI processing: First, based on the specific video processing task, the corresponding AI model is called for processing, and then the output of the AI model is integrated. The AI model includes object detection and recognition, video restoration and enhancement, video content understanding and analysis, video editing and creation, and video content security detection; S4, business logic: Based on business needs and AI-processed data, specific business logic operations are performed; S5, user interaction: providing an operation interface for users; S6, intelligent interaction: Based on user interaction, it can interact with users through voice, gestures, and eye contact to optimize the user's interactive experience.
2. The video processing method based on AI support according to claim 1, characterized in that: The specific steps of target detection and recognition are as follows: first, call the target detection algorithm model, then input the video frame data into the model so that the model can extract features from the image through a convolutional neural network. At the same time, the region generation network is used to generate candidate regions that may contain targets. Then, through classification and regression operations, the targets in the video frames are identified and their categories and location information are determined.
3. The video processing method based on AI support according to claim 1, characterized in that: The specific steps of the video restoration and enhancement are: first call the generative adversarial network model, then use the generator to receive the damaged video frame image as input, and try to generate the restored image by learning the features of a large number of high-quality images. Then the discriminator distinguishes the generated image from the real high-quality image. The two are trained against each other, and the parameters of the generator are continuously iterated to make the generated restored image close to the real high-quality image, thereby realizing the restoration and enhancement of video quality.
4. The video processing method based on AI support according to claim 1, characterized in that: The specific steps of video content understanding and analysis are as follows: first, use computer vision technology such as image segmentation and target tracking to identify and analyze the characters, objects, and scene elements in the video; at the same time, perform speech recognition on the audio in the video and convert it into text; then use natural language processing technology to perform sentiment analysis and topic classification on the text; then, fuse the visual and language information to understand the semantic information in the video, and realize the classification, sentiment analysis, and behavior analysis of the video content.
5. The video processing method based on AI support according to claim 1, characterized in that: The specific steps of video editing and creation are: first call the intelligent algorithm model, then automatically screen relevant materials according to the video theme and style requirements set by the user, and then edit and splice the screened materials through time series analysis and editing algorithms, and add corresponding special effects, subtitles and music.
6. The AI-supported video processing method according to claim 1, characterized in that: The specific steps of the video content security detection are: first calling the multimodal fusion deep learning model, and then analyzing the video image, audio, and text information to analyze whether there is any bad content in the video.
7. The AI-supported video processing method according to claim 1, characterized in that: The specific steps of S4 are: S41, receiving and processing: receiving data from AI processing and integration; S42, Business Requirements Analysis: Based on the business requests submitted by users during user interaction or the business rules preset by the system, the business requirements are analyzed in depth to clarify the specific goals, requirements and constraints of the current business tasks; S43, Data Processing and Integration: Further data processing and integration of received AI-processed data based on business needs; S44, business logic execution: Execute specific business logic operations based on the analyzed business requirements and the processed and integrated data. The business logic operations include security monitoring services, video editing services, video content review services, and business data analysis services. S45, result packaging and output: First, the results after the business logic execution are packaged and converted into a format that can be recognized and displayed by the user interaction. Then, the packaged results are output to the user interaction for the user to view, download or further operate.
8. The AI-supported video processing method according to claim 7, characterized in that: The specific steps of the security monitoring business are as follows: when the AI processing detects that there are abnormal behavior alarm rules in the video, the business logic triggers the alarm mechanism, sends an alarm message to the relevant personnel, and stores the video clip containing the abnormal behavior in a designated storage location for subsequent review and tracing; The specific steps of the video editing business are: synthesizing the integrated video materials, special effects, subtitles and music content according to the set logic and sequence to generate the final video work; The specific steps of the video content review service are: review and process suspected inappropriate content detected by AI processing according to the criteria for determining inappropriate content; for videos determined to be inappropriate content, remove them from shelves, block them, and delete them according to the prescribed process, and record relevant processing logs; for content that cannot be determined, submit it for manual review; The specific steps of the commercial data analysis business are: conducting commercial data analysis on the analysis results of video data based on AI processing, calculating key indicators, generating data analysis reports, and providing data support for merchants' decision-making.
9. The AI-supported video processing method according to claim 1, characterized in that: The specific steps of S5 are: S51, user input: the user submits an operation request through various terminals in user interaction; S52, request parsing: After receiving user input, the request can be parsed to understand the user's specific needs; S53, interface display preparation: while waiting for the business logic to process the request, the interface can be prepared for display and the interface display content can be updated according to the request type; S54, result reception and analysis: After the business logic completes processing and outputs the result to the user interface, the user interface can first receive the processing result, then parse the result to determine the type and content of the result, and then prepare the corresponding display method according to the different result types; S55, results display: S551, video result display: If the processing result is a video file, a video playback window is provided on the interface, supporting various playback control functions. At the same time, a download button is provided to facilitate users to save the video locally, and a sharing function is provided to enable users to share the video to social media platforms with one click; S552, data result display: The data results of the data analysis report are displayed on the interface in the form of intuitive charts and tables, so that users can quickly understand the meaning of the data.
10. The AI-supported video processing method according to claim 1, characterized in that: The specific steps of S6 are: S61, User Input Capture: Utilizes multiple sensors and input devices to capture user input information in real time; S62, input information preprocessing: preprocessing the captured original input information to improve the accuracy and efficiency of subsequent processing; Perform noise reduction processing on voice data to remove environmental noise interference and enhance the clarity of voice signals; Endpoint detection technology is used to determine the start and end points of speech and segment valid speech segments. Image enhancement and feature extraction are performed on gesture and eye contact interaction data. S63, Command Parsing and Intent Recognition: Utilizes natural language processing and computer vision technologies to deeply analyze pre-processed input information. For voice command processing, after converting speech to text, the system uses a semantic analysis model to understand the user's intended meaning and identify the user's action. For gesture and eye contact operations, the system determines the user's intended action based on pre-defined gesture rules and eye contact interaction logic. S64, Intelligent Response Generation: Generates personalized intelligent responses based on identified user intent, combined with historical user behavior data and preference models. At the same time, it incorporates affective computing technology into the response process, adjusting the response method and content based on the user's current emotional state, providing more emotionally sensitive feedback. S65, interactive feedback execution: Feedback the generated intelligent response to the user in an appropriate manner to achieve human-computer interaction.