PPT content generation and interaction optimization system based on AI drive
Through the AI-based PPT content generation and interaction optimization system, the problem of the lack of real-time interactivity of the existing PPT demonstration mode is solved, efficient content generation and real-time interaction are achieved, and demonstration effect and user experience are improved.
Patent Information
- Application Number
- CN202510354666.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing PPT demonstration mode lacks real-time interactiveness, making it difficult to adjust content based on the audience's real-time concerns and reactions, affecting the effect of information transmission.
Using an AI-driven PPT content generation and interaction optimization system, the AB distinguishes module, content generation module and interaction optimization module are used to realize the automated generation and real-time interaction of PPT content.
It improves the efficiency and quality of PPT production, enhances the interactivity and user experience during the demonstration process, and allows speakers to flexibly adjust content based on the audience's reactions.
Smart Images

Figure CN120012940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimedia technology, and in particular to an AI-driven PPT content generation and interaction optimization system. Background Art
[0002] In today's information dissemination and presentation scenarios, PPT is a widely used presentation tool.
[0003] However, the existing PPT presentation mode has obvious drawbacks. The presentation process basically proceeds page by page in the order set by the speaker in advance. The audience is in a state of passively receiving information and lacks a way to actively participate. It is difficult for the speaker to flexibly adjust the PPT presentation content according to the audience's real-time focus, interest points and immediate reactions, which to a certain extent affects the effect of information communication and the interactivity of the presentation.
[0004] With the continuous development of intelligent interactive technologies such as artificial intelligence and computer vision, a system is needed that can more conveniently generate PPT content and enable PPT to better interact with the audience during the presentation process to address the above-mentioned issues. Summary of the invention
[0005] The purpose of the present invention is to propose an AI-driven PPT content generation and interactive optimization system in order to solve the above-mentioned problems.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] The AI-driven PPT content generation and interactive optimization system includes the following parts:
[0008] AB division module: divides the PPT page into area A and area B, and determines the display content categories corresponding to area A and area B. The proportion of area A and area B on the page can be customized by the user according to their own needs;
[0009] Content generation module: Based on the PPT theme and text entered by the user, the content of Area A and Area B of each PPT page is filled through analysis using natural language processing technology;
[0010] Interaction optimization module: After the PPT presenter selects any person in the audience as the target user, the target user's line of sight information is collected and analyzed to determine the area on the PPT corresponding to the line of sight; the target user's hand movements are collected and analyzed to identify gestures, and the identified gestures are matched with the gestures corresponding to the preset operation instructions;
[0011] Storage module: uses a combination of relational database and non-relational database to store data.
[0012] Preferably, the AB distinction module also includes the following contents:
[0013] Each PPT page is divided into area A and area B, and area A and area B are each further divided into a preset number of areas; when the user's eyes stay on any area for a preset time, the area of the area will be enlarged according to a preset ratio, and a preset number of gesture category symbols will pop up around the area, and each gesture category symbol corresponds to an operation instruction.
[0014] Preferably, the content generation module specifically includes the following contents:
[0015] After the user inputs the PPT topic and text, AI language processing technology is used to analyze the content. The analysis results are obtained through lexical analysis, syntactic analysis, and semantic analysis. The content of area A and area B is filled in according to the pre-set rules.
[0016] And according to the quantity and type of the contents in area A and area B, the layout adjustment algorithm is automatically started.
[0017] Preferably, the content generation module further includes the following contents:
[0018] According to the different contents of area A and area B, select and match corresponding elements from the preset element library;
[0019] When there is data content in area A, appropriate charts are automatically generated based on data characteristics and display requirements;
[0020] For the case descriptions in Area B, search for matching pictures or icons from the picture library based on the case theme and key information;
[0021] After the PPT is initially generated, the user can perform corresponding operations through the operation interface.
[0022] Preferably, the collecting and analyzing the sight line information of the target user to determine the area corresponding to the sight line on the PPT specifically includes the following parts:
[0023] The image information of the sight direction data corresponding to the target user information is obtained, and the image information taken at each time interval is screened to retain the image containing the complete eye area of the target user; the image is grayed, and a threshold is set according to the grayscale difference between the pupil and the surrounding area to segment the pupil; the pupil edge is detected by combining the edge detection algorithm, and then the pupil center coordinates are determined by ellipse fitting;
[0024] The direction of the target user's sight line in the real space is determined through analysis, and the PPT area corresponding to the target user's sight line is determined according to the direction of the target user's sight line in the real space.
[0025] Preferably, the collecting and analyzing the hand movements of the target user to identify the gestures specifically includes the following parts:
[0026] The hand images of the target user are collected at preset time intervals, and the hand images are preprocessed to determine the hand contour edges in the images, and the edge images are contour tracked to connect the edge points to form a closed contour; wherein the hand images include the hand images holding the microphone and the hand images making gestures, and the hand images holding the microphone and the hand images making gestures are analyzed to keep the hand images making gestures and mark them as key images, and the hand holding the microphone and the key images are processed accordingly at the same time, wherein the key images are processed as follows:
[0027] The area of the contour region enclosed by the contour in the key image is calculated by Green's formula;
[0028] The distances between all edge points on the contour are summed up to calculate the contour perimeter;
[0029] After obtaining the contour area and perimeter as feature vectors, similarity calculation is performed with the feature vectors corresponding to each gesture in the preset gesture library in turn, wherein the similarity calculation includes Euclidean distance calculation and cosine similarity calculation of the feature vectors to obtain feature distance value and cosine evaluation value;
[0030] Calculate the absolute value of the difference between the cosine similarity and 1 to obtain the cosine evaluation value;
[0031] Preset the weight factors of the feature distance value and the cosine evaluation value, respectively multiply the feature distance value and the cosine evaluation value with their corresponding weight factors, and then perform sum calculation to obtain the matching deviation coefficient;
[0032] The matching deviation coefficients of the collected gestures and each gesture in the preset gesture library are calculated in turn, and a matching deviation coefficient threshold is preset. Each matching deviation coefficient obtained is compared with the matching deviation coefficient threshold, and the gesture corresponding to the matching deviation coefficient lower than the matching deviation coefficient threshold is determined as a gesture category.
[0033] Preferably, the acquisition of the key image specifically includes the following parts:
[0034] Detect a preset number of key points of the hand through the hand key point detection model, and obtain the coordinates of the hand key points in the image coordinate system through the detection algorithm of the collected image;
[0035] Calculate the displacement of each key point in the adjacent images collected on the timeline, preset a displacement threshold, compare the displacement of each key point in each adjacent image with the displacement threshold, mark the key points with a displacement less than the displacement threshold as stay points, obtain the number of stay points in all adjacent images, and take any image of the adjacent images with the largest number of stay points as the key image.
[0036] Preferably, the hand holding the microphone is analyzed as follows:
[0037] When the target user holds the microphone, the pressure value obtained by the sensor on the microphone changes, and after the pressure value change amplitude exceeds a preset amplitude range and remains for a preset time, the target user's hand holding pressure and heart rate are obtained at preset time intervals;
[0038] The grip pressure values obtained at each time point are summed and divided by the number of grip pressure values to obtain the pressure mean;
[0039] Arrange the grip pressure values from left to right in the order of acquisition time, calculate the difference between the next grip pressure value and the previous grip pressure value in the time series, and take the absolute value to obtain the pressure difference;
[0040] Extract the maximum grip pressure value and the minimum grip pressure value from each grip pressure value, perform difference calculation on them, and take the absolute value to obtain the extreme pressure value;
[0041] Preset weight factors for the pressure mean, pressure difference and pressure extreme value, respectively multiply the pressure mean, pressure difference and pressure extreme value with their corresponding weight factors and then sum them up to obtain a comprehensive pressure value;
[0042] According to the above process of analyzing the grip pressure to obtain the pressure comprehensive value, the heartbeat frequency is analyzed to obtain the heartbeat comprehensive value;
[0043] The weight factors of the pressure comprehensive value and the heartbeat comprehensive value are preset, and the pressure comprehensive value and the heartbeat comprehensive value are respectively multiplied by their corresponding weight factors and then summed to obtain a pressure coefficient, and corresponding processing is performed according to the pressure coefficient.
[0044] Preferably, the corresponding processing according to the pressure coefficient specifically includes:
[0045] Three groups of threshold value ranges are preset, and each group of threshold value ranges corresponds to a stress level. The pressure coefficient is matched with the value ranges of the three thresholds to obtain the stress level corresponding to the pressure coefficient, where the stress levels include: normal, mild, and abnormal.
[0046] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0047] 1. The present invention automatically classifies important content into area A and auxiliary information into area B based on the PPT theme and text input by the user and through AI language processing technology analysis. It can also automatically start the layout adjustment algorithm according to the quantity and type of content in areas A and B, expand or reduce the proportion of the corresponding area, adopt appropriate typesetting methods, and follow design principles to achieve the best visual presentation effect, greatly improving the efficiency and quality of PPT production.
[0048] 2. The present invention realizes the execution of corresponding operation instructions of PPT by collecting and analyzing the line of sight and hand movements of the target user, and can analyze the tension level of the target user according to the pressure value, heart rate, etc. when holding the microphone, and display the tension level on the PPT screen with different color symbols, which is convenient for the presenter to take corresponding guiding measures, greatly improving the interactivity and user experience in the PPT presentation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0050] Figure 1 is a flow chart of the present invention; DETAILED DESCRIPTION
[0051] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present application. The present application can be embodied in many different forms and purposes and should not be limited to the embodiments described herein. These embodiments are provided to make the present application comprehensive and complete, and to fully convey the scope of the present application to those skilled in the art. The embodiments do not limit the present application.
[0052] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and / or this specification, and will not be interpreted in an idealized or overly formal sense, unless explicitly defined as such herein.
[0053] See also Figure 1 As shown, the present invention provides a technical solution:
[0054] The AI-driven PPT content generation and interactive optimization system includes the following parts:
[0055] AB division module: divides the PPT page into area A and area B, and determines the display content categories corresponding to area A and area B. The proportion of area A and area B on the page can be customized by the user according to their own needs;
[0056] Content generation module: Based on the PPT theme and text entered by the user, the content of Area A and Area B of each PPT page is filled through analysis using natural language processing technology;
[0057] Content generation module, specifically including the following contents:
[0058] After the user inputs the PPT topic and text, AI language processing technology is used to analyze the content. Through lexical analysis, the keywords, nouns, verbs, adjectives and other vocabulary categories in the text are accurately identified; then with the help of syntactic analysis, the grammatical structure of the sentence is clarified, and the subject, predicate, object, attributive, adverbial, complement and other components are determined; finally, through semantic analysis, the meaning of the text expression, the theme and the logical connection between the various parts are deeply understood; based on the analysis results, the content of area A and area B is filled in according to the pre-set rules;
[0059] For example, important contents such as core ideas and main arguments can be classified into Area A, and auxiliary information such as case descriptions, supplementary data, and references can be classified into Area B. For example, in a business report PPT, important contents such as business growth data and key decisions can be placed in Area A, while auxiliary contents such as market research cases and brief analysis of competitors can be placed in Area B.
[0060] And according to the quantity and type of content in area A and area B, the layout adjustment algorithm is automatically started;
[0061] If there is a lot of text in area A, the system will appropriately expand the proportion of area A and reduce area B, and adopt appropriate text layout methods, such as adjusting line spacing, font size, paragraph indentation, etc., to ensure that the content of area A is complete and clearly displayed; if there are many pictures or charts added to area B, the system will increase the space of area B and arrange these elements reasonably to avoid crowding or incoordination;
[0062] During layout adjustment, follow design principles, such as maintaining alignment and symmetry of page elements, and maintaining the overall coordination and unity of color matching, to achieve the best visual presentation effect;
[0063] Also included are the following:
[0064] According to the different contents of area A and area B, select and match corresponding elements from the preset element library;
[0065] When there is data content in area A, appropriate charts are automatically generated based on data characteristics and display requirements; such as bar charts, line charts, pie charts, scatter charts, etc.
[0066] For example, for time series data, line graphs are preferred to show the trend of change; for percentage data, pie charts are generated for intuitive presentation; for case descriptions in Area B, relevant pictures or icons are retrieved from the picture library based on the case theme and key information to enhance the visualization and appeal of the content and help users better understand the content;
[0067] After the PPT is initially generated, users can perform corresponding operations through the operation interface;
[0068] For example, users can manually adjust the ratio of area A and area B to meet special content display needs; they can change element styles at will, such as changing chart color, line thickness, font style, etc.; they can also easily add or delete content, supplement missing information or delete redundant content. The system will respond in real time and update synchronously. According to each step of the user's operation, it will quickly recalculate the layout, adjust the element matching, and present the modified effect in real time until a PPT document that meets the user's wishes is generated;
[0069] The AB differentiation module also includes the following:
[0070] Each page of the PPT is divided into area A and area B, and each area A and area B is further divided into a preset number of areas; when the user's line of sight stays on any area for a preset time, the area will be enlarged according to a preset ratio, and a preset number of gesture category symbols will pop up around the area, and each gesture category symbol corresponds to an operation instruction;
[0071] Interactive optimization module: After the PPT presenter selects any person in the audience as the target user, the microphone is passed to the target user. When the microphone is positioned within the preset range for a preset time, the camera rotation angle is calculated using trigonometric functions and geometric principles based on the microphone positioning information and the current position and orientation of the camera, so that the camera is rotated to a corresponding angle on the pan / tilt and then aimed at the target user for shooting; the target user's line of sight information is collected and analyzed to determine the area corresponding to the line of sight on the PPT; the target user's hand movements are collected and analyzed to identify gestures, and the identified gestures are matched with the gestures corresponding to the preset operation instructions. Once the match is successful, the corresponding operation instructions are executed on the PPT;
[0072] Collect and analyze the target user's sight information to determine the area on the PPT corresponding to the sight, including the following parts:
[0073] The image information of the sight direction data corresponding to the target user information is obtained, and the image information taken at each time interval is screened to retain the image containing the complete eye area of the target user; the image is grayed, and a threshold is set according to the grayscale difference between the pupil and the surrounding area to segment the pupil; the pupil edge is detected by combining the edge detection algorithm, and then the pupil center coordinates are determined by ellipse fitting;
[0074] The image contrast is enhanced by histogram equalization to highlight the corneal reflection point; the edge detection algorithm is used to detect the edge information in the image and preliminarily locate the position of the reflection point; according to the grayscale characteristics of the corneal reflection point, a threshold is set, and the pixel points in the image with grayscale values higher than the threshold are regarded as possible corneal reflection points;
[0075] Create a template with similar shape and features to the possible corneal reflection point, and perform template matching in the image; calculate the similarity between the template and each area in the image, find the area with the highest similarity, and determine the possible corneal reflection point as the corneal reflection point;
[0076] The calculation of the similarity between the template and each area in the image includes the following parts:
[0077] Normalize the template image T(x, y) and the image to be matched I(x, y) to map their grayscale values to the range of 0 to 1; then for each position (i, j) in the image to be matched, calculate the correlation coefficient between the template and the area of the same size centered at (i, j); the calculation formula is in is the average gray value of the template image, It is the average gray value of the area centered at (i, j); the value range of NCC is between -1 and 1. The closer the value is to 1, the higher the similarity between the template and the area. By traversing all positions of the image to be matched, the position of the maximum NCC value is determined as the best matching position of the corneal reflection point;
[0078] Approximate the eyeball as a sphere, and take the eyeball radius, pupil center, corneal reflection point, and internal and external parameters of the camera; the internal parameters include focal length and principal point coordinates, and the external parameters include rotation matrix and translation vector; calculate the coordinates of the world coordinate system through the formula to determine the direction of the target user's line of sight in the real space; and determine the PPT area corresponding to the target user's line of sight based on the direction of the target user's line of sight in the real space;
[0079] The specific calculation process includes the following parts:
[0080] The eyeball is approximately regarded as a sphere. Let the radius of the sphere be r. Connect the position of the corneal reflection point R with the pupil center P to form a vector
[0081] Get the focal length f of the camera and the coordinates of the principal point (c x ,c y ), forming the internal parameter matrix where f x and f y are the focal lengths in the x and y directions respectively;
[0082] The rotation matrix R is established based on the external parameters of the camera to describe the rotation relationship between the camera and the world coordinate system. The translation vector T describes the position of the camera origin in the world coordinate system. Together, they form the external parameter matrix [R|T], which is used to convert the coordinates of the camera coordinate system and the coordinates of the world coordinate system.
[0083] Obtain the coordinates of P and R in the image plane (u p ,v p )、(u R ,v R ), and use the internal parameter matrix K to convert it into the camera coordinate system coordinates (X cP ,Y cP ,Z cP )、(X cR ,Y cR ,Z cR ), and then use the external parameter matrix to convert it into the world coordinate system coordinates (X wP ,Y wP ,Z wP )、(X wR ,Y wR ,Z wR ) calculates the vector Normalization process to get unit vector The direction of the line of sight in real space can be determined;
[0084] Collect and analyze the target user's hand movements to identify gestures, which includes the following parts:
[0085] The hand images of the target user are collected at preset time intervals, and the hand images are preprocessed to determine the hand contour edges in the images, and the edge images are contour tracked to connect the edge points to form a closed contour; wherein the hand images include the hand images holding the microphone and the hand images making gestures, and the hand images holding the microphone and the hand images making gestures are analyzed to keep the hand images making gestures and mark them as key images, and the hand holding the microphone and the key images are processed accordingly at the same time, wherein the key images are processed as follows:
[0086] Analyze the collected images to retain the key images, including the following parts:
[0087] Detect a preset number of key points of the hand through the hand key point detection model, and obtain the coordinates of the hand key points in the image coordinate system through the detection algorithm of the collected image;
[0088] Calculate the displacement of each key point in the adjacent images collected on the timeline, preset a displacement threshold, compare the displacement of each key point in each adjacent image with the displacement threshold, mark the key point whose displacement is less than the displacement threshold as a stay point, obtain the number of stay points in all adjacent images, and take any image of the adjacent images with the largest number of stay points as a key image;
[0089] Preprocessing includes graying the key image, denoising, and detecting the image contour edge using an edge detection algorithm;
[0090] The area of the contour region enclosed by the contour in the key image is calculated by Green's formula;
[0091] The distances between all edge points on the contour are summed up to calculate the contour perimeter;
[0092] For the adjacent points P on the contour i (x i ,y i ) and P i+1 (x i+1 ,y i+1 ), the distance between the two points is
[0093]
[0094] The perimeter is: Where n is the number of points on the contour;
[0095] After obtaining the contour area and perimeter as feature vectors, similarity calculation is performed with the feature vectors corresponding to each gesture in the preset gesture library in turn, wherein the similarity calculation includes Euclidean distance calculation and cosine similarity calculation of the feature vectors to obtain feature distance value and cosine evaluation value;
[0096] Euclidean distance calculation: Calculate the Euclidean distance between two feature vectors X and Y. The formula is:
[0097] where x i and i The eigenvectors are and The i-th element of , n represents the dimension of the feature vector, that is, the number of elements in the vector; the smaller d is, the more similar the features of the two gestures are;
[0098] Cosine similarity: Calculate the cosine similarity of two feature vectors X and Y. The formula is:
[0099] The closer the sim value is to 1, the more similar the features of the two gestures are;
[0100] Calculate the absolute value of the difference between the cosine similarity and 1 to obtain the cosine evaluation value;
[0101] Preset the weight factors of the feature distance value and the cosine evaluation value, respectively multiply the feature distance value and the cosine evaluation value with their corresponding weight factors, and then perform sum calculation to obtain the matching deviation coefficient;
[0102] Calculate the matching deviation coefficients of the collected gestures and each gesture in the preset gesture library in turn, preset a matching deviation coefficient threshold, compare each obtained matching deviation coefficient with the matching deviation coefficient threshold, and determine the gesture corresponding to the matching deviation coefficient lower than the matching deviation coefficient threshold as a gesture category;
[0103] The hand holding the microphone is analyzed as follows:
[0104] When the target user holds the microphone, the pressure value obtained by the sensor on the microphone changes, and after the pressure value change amplitude exceeds a preset amplitude range and remains for a preset time, the target user's hand holding pressure and heart rate are obtained at preset time intervals;
[0105] The grip pressure values obtained at each time point are summed and divided by the number of grip pressure values to obtain the pressure mean;
[0106] Arrange the grip pressure values from left to right in the order of acquisition time, calculate the difference between the next grip pressure value and the previous grip pressure value in the time series, and take the absolute value to obtain the pressure difference;
[0107] Extract the maximum grip pressure value and the minimum grip pressure value from each grip pressure value, perform difference calculation on them, and take the absolute value to obtain the extreme pressure value;
[0108] Preset weight factors for the pressure mean, pressure difference and pressure extreme value, respectively multiply the pressure mean, pressure difference and pressure extreme value with their corresponding weight factors and then sum them up to obtain a comprehensive pressure value;
[0109] According to the above process of analyzing the grip pressure to obtain the pressure comprehensive value, the heartbeat frequency is analyzed to obtain the heartbeat comprehensive value;
[0110] The weight factors of the pressure comprehensive value and the heartbeat comprehensive value are preset, and the pressure comprehensive value and the heartbeat comprehensive value are respectively multiplied by the corresponding weight factors and then summed to obtain the pressure coefficient, and corresponding processing is performed according to the pressure coefficient;
[0111] According to the pressure coefficient, corresponding processing is carried out, including:
[0112] Three groups of threshold value ranges are preset, and each group of threshold value ranges corresponds to a stress level. The pressure coefficient is matched with the value ranges of the three thresholds to obtain the stress level corresponding to the pressure coefficient, where the stress levels include: normal, mild, and abnormal;
[0113] The tension level is represented by symbols, and the color of the symbols for each tension level is different; when the tension level is normal, mild, and abnormal, the corresponding symbol colors are green, yellow, and red respectively;
[0114] The symbols are displayed on the PPT display screen. The PPT presenter judges the target user's nervousness based on the symbols and colors corresponding to the nervousness level, and provides corresponding guidance to the target user to relieve his or her nervousness.
[0115] For example:
[0116] When the tension level is mild (yellow triangle icon), the presenter can ease the target user's tension by slowing down the speech speed, increasing the number of interactive question sessions, etc.
[0117] When the tension level is abnormal (red square icon), the presenter can pause the demonstration and give a brief psychological soothing speech, such as "Please relax, the following content will gradually become clear", and combine it with simple relaxation guidance, such as deep breathing exercises, to relieve the tension of the target users;
[0118] Storage module: Data is stored in a combination of relational database and non-relational database. Relational database is used to store structured data, such as user operation logs, and follows SQL query language specifications. Non-relational database is used to store unstructured data, such as material files, and is stored in BSON format.
[0119] The above formulas are obtained by collecting a large amount of data and performing software simulation, and a formula close to the actual value is selected. The influencing weight factor and specific coefficient value in the formula are set by technical personnel in this field according to actual conditions, and can be adjusted and modified later.
[0120] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. AI-driven PPT content generation and interactive optimization system, characterized by: Includes the following parts: AB division module: divides the PPT page into area A and area B, and determines the display content categories corresponding to area A and area B. The proportion of area A and area B on the page can be customized by the user according to their own needs; Content generation module: Based on the PPT theme and text entered by the user, the content of Area A and Area B of each PPT page is filled through analysis using natural language processing technology; Interaction optimization module: After the PPT presenter selects any person in the audience as the target user, the target user's sight information is collected and analyzed to determine the area on the PPT corresponding to the sight line; Collect and analyze the hand movements of the target user to identify the gestures, and match the identified gestures with the gestures corresponding to the preset operation instructions; Storage module: uses a combination of relational database and non-relational database to store data.
2. The AI-driven PPT content generation and interactive optimization system according to claim 1 is characterized in that: The AB distinction module also includes the following contents: Each PPT page is divided into area A and area B, and area A and area B are each further divided into a preset number of areas; when the user's eyes stay on any area for a preset time, the area of the area will be enlarged according to a preset ratio, and a preset number of gesture category symbols will pop up around the area, and each gesture category symbol corresponds to an operation instruction.
3. The AI-driven PPT content generation and interactive optimization system according to claim 1 is characterized in that: The content generation module specifically includes the following contents: After the user inputs the PPT topic and text, AI language processing technology is used to analyze the content. The analysis results are obtained through lexical analysis, syntactic analysis, and semantic analysis. The content of area A and area B is filled in according to the pre-set rules. And according to the quantity and type of the contents in area A and area B, the layout adjustment algorithm is automatically started.
4. The AI-driven PPT content generation and interactive optimization system according to claim 3 is characterized in that: The content generation module further includes the following contents: According to the different contents of area A and area B, select and match corresponding elements from the preset element library; When there is data content in area A, appropriate charts are automatically generated based on data characteristics and display requirements; For the case descriptions in Area B, search for matching pictures or icons from the picture library based on the case theme and key information; After the PPT is initially generated, the user can perform corresponding operations through the operation interface.
5. The AI-driven PPT content generation and interactive optimization system according to claim 1 is characterized in that: The collecting and analyzing of the sight line information of the target user to determine the area corresponding to the sight line on the PPT specifically includes the following parts: The image information of the sight direction data corresponding to the target user information is obtained, and the image information taken at each time interval is filtered to retain the image containing the complete eye area of the target user; The image is grayed out, and a threshold is set according to the grayscale difference between the pupil and the surrounding area to segment the pupil. The pupil edge is detected by combining edge detection algorithm, and then the pupil center coordinates are determined by ellipse fitting; The direction of the target user's sight line in the real space is determined through analysis, and the PPT area corresponding to the target user's sight line is determined according to the direction of the target user's sight line in the real space.
6. The AI-driven PPT content generation and interactive optimization system according to claim 1 is characterized in that: The collecting and analyzing of the hand movements of the target user to identify the gestures specifically includes the following parts: The hand images of the target user are collected at preset time intervals, and the hand images are preprocessed to determine the hand contour edges in the images, and the edge images are contour tracked to connect the edge points to form a closed contour; wherein the hand images include the hand images holding the microphone and the hand images making gestures, and the hand images holding the microphone and the hand images making gestures are analyzed to keep the hand images making gestures and mark them as key images, and the hand holding the microphone and the key images are processed accordingly at the same time, wherein the key images are processed as follows: The area of the contour region enclosed by the contour in the key image is calculated by Green's formula; The distances between all edge points on the contour are summed up to calculate the contour perimeter; After obtaining the contour area and perimeter as feature vectors, similarity calculation is performed with the feature vectors corresponding to each gesture in the preset gesture library in turn, wherein the similarity calculation includes Euclidean distance calculation and cosine similarity calculation of the feature vectors to obtain feature distance value and cosine evaluation value; Calculate the absolute value of the difference between the cosine similarity and 1 to obtain the cosine evaluation value; Preset the weight factors of the feature distance value and the cosine evaluation value, respectively multiply the feature distance value and the cosine evaluation value with their corresponding weight factors, and then perform sum calculation to obtain the matching deviation coefficient; The matching deviation coefficients of the collected gestures and each gesture in the preset gesture library are calculated in turn, and a matching deviation coefficient threshold is preset. Each matching deviation coefficient obtained is compared with the matching deviation coefficient threshold, and the gesture corresponding to the matching deviation coefficient lower than the matching deviation coefficient threshold is determined as a gesture category.
7. The AI-driven PPT content generation and interactive optimization system according to claim 6 is characterized in that: The acquisition of key images specifically includes the following parts: Detect a preset number of key points of the hand through the hand key point detection model, and obtain the coordinates of the hand key points in the image coordinate system through the detection algorithm of the collected image; Calculate the displacement of each key point in the adjacent images collected on the timeline, preset a displacement threshold, compare the displacement of each key point in each adjacent image with the displacement threshold, mark the key points with a displacement less than the displacement threshold as stay points, obtain the number of stay points in all adjacent images, and take any image of the adjacent images with the largest number of stay points as the key image.
8. The AI-driven PPT content generation and interactive optimization system according to claim 6 is characterized in that: The hand holding the microphone is analyzed as follows: When the target user holds the microphone, the pressure value obtained by the sensor on the microphone changes, and after the pressure value change amplitude exceeds a preset amplitude range and remains for a preset time, the target user's hand holding pressure and heart rate are obtained at preset time intervals; The grip pressure values obtained at each time point are summed and divided by the number of grip pressure values to obtain the pressure mean; Arrange the grip pressure values from left to right in the order of acquisition time, calculate the difference between the next grip pressure value and the previous grip pressure value in the time series, and take the absolute value to obtain the pressure difference; Extract the maximum grip pressure value and the minimum grip pressure value from each grip pressure value, perform difference calculation on them, and take the absolute value to obtain the extreme pressure value; Preset weight factors for the pressure mean, pressure difference and pressure extreme value, respectively multiply the pressure mean, pressure difference and pressure extreme value with their corresponding weight factors and then sum them up to obtain a comprehensive pressure value; According to the above process of analyzing the grip pressure to obtain the pressure comprehensive value, the heartbeat frequency is analyzed to obtain the heartbeat comprehensive value; The weight factors of the pressure comprehensive value and the heartbeat comprehensive value are preset, and the pressure comprehensive value and the heartbeat comprehensive value are respectively multiplied by their corresponding weight factors and then summed to obtain a pressure coefficient, and corresponding processing is performed according to the pressure coefficient.
9. The AI-driven PPT content generation and interactive optimization system according to claim 8, characterized in that: The corresponding processing according to the pressure coefficient specifically includes: Three groups of threshold value ranges are preset, and each group of threshold value ranges corresponds to a stress level. The pressure coefficient is matched with the value ranges of the three thresholds to obtain the stress level corresponding to the pressure coefficient, where the stress levels include: normal, mild, and abnormal.
Citation Information
Patent Citations
PPT demonstration assisting system based on Kinect
CN106125928A
Driving auxiliary method and system for monitoring mental status of driver
CN110403617A
Driver behavior state data acquisition device and detection method
CN113094930A
Control method of handheld emergency stop device
CN113655845A
PPT automatic generation method based on AI
CN116975329A