An AI emotional interaction guiding system for children's picture book reading

By constructing an attention distribution plane and optimizing the emotional state analysis model, the visual content of picture books is dynamically adjusted, solving the problem that existing systems cannot accurately capture children's attention distribution, and realizing precise identification and personalized guidance of children's reading status.

CN121807207BActive Publication Date: 2026-05-15XIAMEN SANDU EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN SANDU EDUCATION TECH CO LTD
Filing Date
2026-03-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing interactive systems for children's picture book reading cannot accurately capture children's attention distribution information, resulting in an inability to specifically guide children to maintain focus. The existing guidance methods are also limited and cannot adapt to children's specific reading states.

Method used

By collecting reading behavior data streams, we construct an attention distribution plane, generate refined attention distribution units, optimize the emotional state analysis model, and dynamically adjust the visual content of picture books to achieve emotional interaction guidance.

Benefits of technology

It accurately captures the areas and density of children's attention during reading, improves the accuracy of attention recognition, and enables intelligent, precise identification and personalized guidance of children's reading status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807207B_ABST
    Figure CN121807207B_ABST
Patent Text Reader

Abstract

The application provides an AI emotional interaction guiding system for children's picture book reading, and relates to the technical field of data processing, which comprises: a part for generating the basic visual composition and frame layout of a picture book page according to an AI picture book strategy parameter set; a part for determining the position, angle and form adjustment rules of each component of the basic visual composition based on the emotional guiding parameters in the AI picture book strategy parameter set; a part for generating the picture book visual content by moving the position, rotating the angle and adapting the form of the basic visual composition and frame layout according to the form adjustment rules; a part for collecting the dynamic visual attention data stream of the reader in real time during the presentation process of the picture book visual content, converting the dynamic visual attention data stream into updated cognitive distribution characteristic values, feeding back the updated cognitive distribution characteristic values to an emotional state analysis model, dynamically adjusting the generation of subsequent picture book visual content, and realizing emotional interaction guiding. The application effectively improves the reading concentration and reading experience of children.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an AI-powered emotional interaction guidance system for children's picture book reading. Background Technology

[0002] With the widespread application of interactive children's picture books on smart devices such as tablets and smart learning terminals, using intelligent technology to help children maintain reading focus and optimize the reading interaction experience has become an important development direction in the field of children's intelligent reading. Currently, some picture book reading interaction systems have basic eye-tracking functions, which can collect children's eye information while reading through the device's built-in camera. Its core function is only to roughly determine whether the child is looking at the picture book screen interface, which may be difficult to capture more refined details of attention distribution. When the system detects that the child's gaze has shifted or their attention has been distracted (such as looking to the outside of the screen or looking at a blank area of ​​the screen), the existing guidance methods are relatively simple and fixed. They usually involve changing the overall page color style (such as changing a light-colored page to a more vibrant warm color), playing a preset prompt animation in a fixed corner of the screen (such as a continuously flashing small icon), or following the prompts. The system plays a fixed, encouraging voice message (such as "Please focus on reading") to try and guide children back to reading. While this method serves as a basic reading reminder, there is significant room for improvement. For example, when a child is reading a page of a picture book, their gaze may linger on the cartoon character in the center of the page for a long time, or frequently move between the cartoon character and the short text below. The existing system cannot analyze this kind of fine-grained attention distribution information. It may not be able to identify that the child is currently most focused on the cartoon character, nor can it capture the movement of the gaze between the character and the text. If the child's attention is slightly distracted at this time, the system will still trigger the fixed guidance method mentioned above. It cannot specifically move the cartoon character that the child is focusing on to a position that is easy for the child to catch their gaze, fine-tune the character's orientation to better match the child's focus, or fine-tune the character's proportions to enhance its attractiveness. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide an AI-powered emotional interaction guidance system for children's picture book reading, which effectively improves children's reading focus and reading experience.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] Firstly, an AI-powered emotional interaction guidance system for children's picture book reading includes:

[0006] The data acquisition module is used to collect data streams of children's reading behavior.

[0007] The calculation module is used to construct an attention distribution plane based on the reading behavior data stream; based on the attention distribution plane, it identifies the areas where children's attention is concentrated and calculates the minimum bounding rectangle of the concentrated areas to generate multiple attention distribution units; based on the distribution density and persistence intensity of the reading behavior data stream in each attention distribution unit, it calculates the cognitive distribution feature value.

[0008] The optimization module is used to optimize the emotional state analysis model based on cognitive distribution feature values. The optimized model is then used to process the reading behavior data stream to obtain structured analysis results that characterize children's cognitive and emotional states.

[0009] The matching module is used to perform matching and mapping based on the structured analysis results to obtain a set of AI picture book strategy parameters;

[0010] The processing module is used to generate the basic visual composition and framework layout of the picture book page based on the AI ​​picture book strategy parameter set; determine the position, angle and shape adjustment rules of each component of the basic visual composition based on the emotional guidance parameters in the AI ​​picture book strategy parameter set; and perform position movement, angle rotation and shape adaptation on the basic visual composition and framework layout according to the shape adjustment rules to generate picture book visual content.

[0011] The update module is used to collect the dynamic visual attention data stream of readers in real time during the presentation of picture book visual content, and convert it into updated cognitive distribution feature values, which are then fed back to the emotional state analysis model to dynamically adjust the generation of subsequent picture book visual content and achieve emotional interaction guidance.

[0012] In a second aspect, a computer-readable storage medium storing a program that, when executed by a processor, implements the system.

[0013] The above-described solution of the present invention has at least the following beneficial effects:

[0014] By collecting reading behavior data streams, constructing attention distribution planes, and generating refined attention distribution units, we can accurately capture the areas, density, and intensity of children's attention during reading, enabling fine-grained quantitative analysis of children's reading attention and improving the accuracy and detail of attention recognition. Based on the cognitive distribution feature values, we can optimize the emotional state analysis model, deeply mine and output structured results of children's cognitive level and emotional state during reading, rather than simply judging whether attention is scattered, thus achieving intelligent and accurate recognition of children's reading state. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of an AI-powered emotional interaction guidance system for children's picture book reading, provided by an embodiment of the present invention.

[0016] Figure 2 This is a flowchart illustrating the process of obtaining a set of AI picture book strategy parameters based on structured analysis results through matching and mapping, as provided in an embodiment of the present invention. Detailed Implementation

[0017] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0018] like Figure 1 As shown, an embodiment of the present invention proposes an AI-powered emotional interaction guidance system for children's picture book reading, comprising:

[0019] The data acquisition module is used to collect data streams of children's reading behavior.

[0020] The calculation module is used to construct an attention distribution plane based on the reading behavior data stream; based on the attention distribution plane, it identifies the areas where children's attention is concentrated and calculates the minimum bounding rectangle of the concentrated areas to generate multiple attention distribution units; based on the distribution density and persistence intensity of the reading behavior data stream in each attention distribution unit, it calculates the cognitive distribution feature value.

[0021] The optimization module is used to optimize the emotional state analysis model based on cognitive distribution feature values. The optimized model is then used to process the reading behavior data stream to obtain structured analysis results that characterize children's cognitive and emotional states.

[0022] The matching module is used to perform matching and mapping based on the structured analysis results to obtain a set of AI picture book strategy parameters;

[0023] The processing module is used to generate the basic visual composition and framework layout of the picture book page based on the AI ​​picture book strategy parameter set; determine the position, angle and shape adjustment rules of each component of the basic visual composition based on the emotional guidance parameters in the AI ​​picture book strategy parameter set; and perform position movement, angle rotation and shape adaptation on the basic visual composition and framework layout according to the shape adjustment rules to generate picture book visual content.

[0024] The update module is used to collect the dynamic visual attention data stream of readers in real time during the presentation of picture book visual content, and convert it into updated cognitive distribution feature values, which are then fed back to the emotional state analysis model to dynamically adjust the generation of subsequent picture book visual content and achieve emotional interaction guidance.

[0025] In this embodiment of the invention, by collecting reading behavior data streams, constructing an attention distribution plane, and generating refined attention distribution units, the attention concentration areas, distribution density, and duration of children's reading can be accurately captured, enabling fine-grained quantitative analysis of children's reading attention and improving the accuracy and detail of attention recognition. Furthermore, by optimizing the emotional state analysis model based on cognitive distribution feature values, it is possible to deeply mine and output structured results of children's cognitive level and emotional state during reading, rather than simply determining whether attention is scattered, thus achieving intelligent and precise recognition of children's reading state.

[0026] In a preferred embodiment of the present invention, the collection of children's reading behavior data stream specifically includes: firstly, collecting children's reading behavior data stream during picture book reading. This data stream originates from a collection device mounted on a smart reading terminal. The collection device specifically consists of a high-definition camera and an eye-tracking device, which work together to form an artificial intelligence data collection front-end. This device captures children's visual gaze information, gaze duration, and other related data in real time. The visual gaze information includes the specific coordinates of the gaze point and the gaze direction. The gaze duration is accurate to the millisecond level. This refined data can provide artificial intelligence with a raw representation of children's attention distribution and reading concentration, avoiding bias in artificial intelligence analysis caused by coarse data. Subsequently, all the captured data is integrated in chronological order to form a complete reading behavior data stream.

[0027] This embodiment, by collecting children's reading behavior data streams in a refined manner, and by specifying the specific type of the collection device and the specific content of the collected data, can accurately analyze the distribution of visual gaze points.

[0028] In a preferred embodiment of the present invention, an attention distribution plane is constructed based on reading behavior data stream; according to the attention distribution plane, the child's attention concentration area is identified, and the minimum bounding rectangle of the concentration area is calculated to generate multiple attention distribution units; based on the distribution density and persistence intensity of the reading behavior data stream in each attention distribution unit, cognitive distribution feature values ​​are calculated, which may include:

[0029] This process involves analyzing the reading behavior data stream, extracting the visual fixation point coordinate sequence, and constructing an attention distribution plane based on the positions of all visual fixations on the picture book page. Specifically, this includes parsing the collected reading behavior data stream using a pre-defined data parsing algorithm. This algorithm is a timestamp-based data stream splitting and parsing algorithm with a preset parsing frequency of 100Hz. The preset data format includes timestamp, x-axis, y-axis, fixation direction, and fixation duration, with values ​​ranging from milliseconds (0~9999999999ms), x-axis 0~1, y-axis 0~1, fixation direction 0~360°, and fixation duration 10~10000ms. The specific calculation process is as follows: First, the continuous reading behavior data stream is split into individual data frames by timestamp according to the preset parsing frequency. Then, the format of each data frame is validated, and data frames that do not conform to the preset format are removed. Finally, the horizontal and vertical coordinates of all visual fixations are extracted from the validated data frames and arranged in the order of the children's gaze time to form a visual fixation coordinate sequence. The preset coordinate system of the picture book page takes the upper left corner of the page as the origin, and the values ​​of the horizontal and vertical coordinates are both in the range of 0 to 1. According to the specific position of each visual fixation point in the preset coordinate system, all visual fixations are mapped one by one to the two-dimensional plane of the picture book page to construct an attention distribution plane that can intuitively reflect the distribution of children's attention.

[0030] On the attention distribution plane, visual fixations are analyzed to identify multiple spatially clustered fixation point groups. The area covered by these fixation point groups is defined as the attention concentration region. Specifically, on the constructed attention distribution plane, the spatial distribution of all visual fixations is analyzed. A preset spatial clustering judgment rule is adopted, which states that if the straight-line distance between two fixations does not exceed 0.05 times the side length of the picture book page, and three or more consecutive fixations meet this distance condition, then these fixations are determined to form a cluster. Based on this, multiple spatially close and clustered fixation point groups are identified. Each fixation point group represents an area where children's attention is relatively concentrated during reading, and the entire spatial range covered by each fixation point group is defined as an attention concentration region.

[0031] Based on the attention-focused region, the gaze coordinates of all visual gaze points constituting the attention-focused region are extracted. Specifically, for each defined attention-focused region, a coordinate extraction algorithm is used. This algorithm is a coordinate traversal and filtering algorithm within the region. The preset filtering threshold is 0.001 units outward from the boundary of the attention-focused region. The specific calculation process is to set a filtering coordinate interval based on the spatial range of the identified attention-focused region. The horizontal coordinate interval is from the minimum horizontal coordinate of the region minus 0.001 to the maximum horizontal coordinate plus 0.001, and the vertical coordinate interval is from the minimum vertical coordinate of the region minus 0.001 to the maximum vertical coordinate plus 0.001. Then, all visual gaze coordinates are traversed, and gaze points whose coordinates fall within the filtering interval are selected as the visual gaze coordinates constituting the attention-focused region.

[0032] Based on all fixation point coordinates, the minimum and maximum values ​​of the x-coordinate and y-coordinate are found. Specifically, this involves: first, arranging all extracted visual fixation point coordinates in order; selecting the x-coordinate of the first fixation point as the initial minimum and maximum x-coordinate values; then, iterating through the x-coordinates of each remaining fixation point, comparing the current fixation point's x-coordinate with the initial minimum x-coordinate value; if the current fixation point's x-coordinate is less than the initial minimum x-coordinate value, updating the initial minimum x-coordinate to the current fixation point's x-coordinate; if the current fixation point's x-coordinate is greater than the initial minimum x-coordinate value... If the value is large, the initial maximum x-coordinate is updated to the x-coordinate of the current fixation point. After traversing all fixation points, the final minimum initial x-coordinate is the minimum x-coordinate of all fixation points, and the final maximum initial x-coordinate is the maximum x-coordinate of all fixation points. The minimum and maximum y-coordinates are obtained in the same way. The y-coordinate of the first fixation point is selected as the initial minimum and maximum y-coordinates. The y-coordinates of each remaining fixation point are traversed one by one, and the initial minimum and maximum y-coordinates are updated after comparison. After the traversal is completed, the minimum and maximum y-coordinates of all fixation points are obtained.

[0033] Based on the minimum x-coordinate and the maximum y-coordinate, determine the coordinates of the top-left vertex of the minimum bounding rectangle; based on the maximum x-coordinate and the minimum y-coordinate, determine the coordinates of the bottom-right vertex of the minimum bounding rectangle. Specifically, the coordinates of the top-left vertex are determined by selecting the minimum x-coordinate of all fixation points as the x-coordinate of the top-left vertex and the maximum y-coordinate of all fixation points as the y-coordinate of the top-left vertex, combining the two to form the coordinates of the top-left vertex of the minimum bounding rectangle, ensuring that this vertex accurately corresponds to the top-left boundary of the attention-focused area; the coordinates of the bottom-right vertex are determined by selecting the maximum x-coordinate of all fixation points as the x-coordinate of the bottom-right vertex and the minimum y-coordinate of all fixation points as the y-coordinate of the bottom-right vertex, combining the two to form the coordinates of the bottom-right vertex of the minimum bounding rectangle.

[0034] Based on the coordinates of the top-left and bottom-right vertices, a minimum bounding rectangle is generated that completely encloses the gaze point coordinates and whose edges are parallel to the coordinate axes of the picture book page. Based on this minimum bounding rectangle, attention distribution units are obtained, specifically including: first, based on the determined coordinates of the top-left and bottom-right vertices, calculating the four sides of the minimum bounding rectangle. The top side is parallel to the x-coordinate of the picture book page, with its x-coordinate ranging from the x-coordinate of the top-left vertex to the x-coordinate of the bottom-right vertex, and its y-coordinate being the y-coordinate of the top-left vertex; the bottom side is parallel to the x-coordinate of the picture book page, with its x-coordinate range the same as the top side, and its y-coordinate being the y-coordinate of the bottom-right vertex; the left side is parallel to the y-coordinate of the picture book page, with its y-coordinate range ranging from the y-coordinate of the bottom-right vertex to the y-coordinate of the top-left vertex, and its x-coordinate being the x-coordinate of the top-left vertex. The x-coordinate of the corner vertex; the y-coordinate of the right side parallel to the picture book page, with the right y-coordinate range being the same as the left; the x-coordinate of the right side being the x-coordinate of the bottom right corner vertex; connect the four sides sequentially to generate the smallest bounding rectangle that can completely enclose all visual fixation points within the attention focus area and whose edges are parallel to the preset coordinate axes of the picture book page. After generation, deviation verification is performed. The preset parallel deviation threshold for the rectangle edge is 0.0001 units in length to ensure that the parallel deviation between the rectangle edge and the coordinate axes does not exceed this threshold. At the same time, it is verified that all fixation points are inside the rectangle, with no omissions and no extra blank areas; each generated smallest bounding rectangle is directly determined as an attention distribution unit. One attention distribution unit is generated for one attention focus area, and multiple attention distribution units are generated for multiple attention focus areas.

[0035] Based on attention distribution units, all visual fixations located within the boundary range of each attention distribution unit are obtained, generating a subset of visual fixations for each attention distribution unit. The duration of each visual fixation within this subset is then obtained, resulting in a duration subset. Specifically, for each attention distribution unit, its boundary range is determined based on the coordinates of its top-left and bottom-right vertices. The horizontal boundary is from the horizontal coordinate of the top-left vertex to the horizontal coordinate of the bottom-right vertex, and the vertical boundary is from the vertical coordinate of the bottom-right vertex to the vertical coordinate of the top-left vertex. Only visual fixations whose horizontal coordinates and vertical coordinates are both within the horizontal and vertical boundary ranges are considered to be within the boundary range of that attention distribution unit. All visual fixations are then iterated through sequentially. Visual fixations are selected based on the boundary conditions. These fixations are then integrated sequentially according to the time of the child's gaze, with secondary verification during the integration process to ensure that each integrated fixation meets the boundary requirements. This results in a subset of visual fixations for this attention distribution unit. Simultaneously, from the original reading behavior data stream, the gaze duration corresponding to each fixation in the visual fixation subset is retrieved based on the timestamp of each fixation. One fixation corresponds to one gaze duration. These durations are then integrated sequentially according to the time of the corresponding fixation. After integration, the durations are checked to ensure that each fixation in the visual fixation subset has a corresponding duration, with no missing or incorrect correspondences. This results in a subset of durations for this attention distribution unit.

[0036] Based on the visual fixation point subset, the total number of visual fixations in the subset is counted. The total number of visual fixations is then divided by the area of ​​the corresponding attention distribution unit to obtain the unit distribution density. Specifically, this involves: using a point-by-point matching counting algorithm based on the visual fixation point subset, with the counting order preset to chronological order of fixation time and a preset allowable deviation value of 0 for secondary verification. The specific calculation process is as follows: first, initialize the count variable to 0; then, traverse each fixation point in the visual fixation point subset one by one according to the chronological order of fixation time. Each time a valid fixation point (i.e., a fixation point whose coordinates conform to the boundary range of the attention distribution unit) is encountered, the count variable is incremented by one. After traversing all fixation points, the value of the count variable is the total number of visual fixations contained in that subset. After the total count is completed, [further steps are taken]. A second check is performed by traversing the subset in reverse and counting each point again. If the two counts are consistent, the count is valid, avoiding counting errors and ensuring the accuracy of the total number of visual fixations. Next, the area of ​​the attention distribution unit is calculated. The side length of the picture book page is preset to 1 unit. Therefore, the area of ​​the attention distribution unit is calculated by multiplying the difference between the maximum and minimum horizontal coordinates of the unit by the difference between the maximum and minimum vertical coordinates of the unit. The total number of visual fixations is then divided by the area of ​​the attention distribution unit. In other words, the total number of visual fixations is used as the dividend and the area of ​​the attention distribution unit is used as the divisor. The result is the unit distribution density of the attention distribution unit.

[0037] Based on a subset of durations, sum all duration values ​​within that subset. Divide the sum by the number of data points in the subset to obtain the unit duration intensity. Specifically, this involves: using a value-by-value summation algorithm with a preset accumulation precision of 0.1ms. The calculation process is as follows: First, initialize the summation variable to 0. Then, according to the time sequence of the corresponding gaze points, iterate through each gaze duration value in the subset, successively adding the currently encountered duration value to the summation variable. That is, first add the first duration value to the summation variable to obtain a new summation variable, then add the second duration value to the new summation variable, and so on, until all duration values ​​have been traversed. The final summation variable is the total duration. One decimal place is retained during the accumulation process to avoid accumulation errors and ensure the accuracy of the total duration summation. Then... The point-to-point matching counting algorithm is used, which is consistent with the point-to-point matching counting algorithm for counting the total number of visual fixations. The preset counting order is the chronological order of the corresponding fixations, and the preset allowable deviation for secondary verification is 0. The specific calculation process is as follows: Initialize the counting variable to 0, and iterate through each data in the duration subset one by one. Each time a valid duration data is encountered, that is, the duration data that corresponds one-to-one with the fixation point in the visual fixation point subset, the counting variable is incremented by one. After the iteration is completed, a reverse secondary verification is performed to ensure that the counting result is accurate. This counting result is the number of data in the duration subset, and this number is completely consistent with the total number of fixations in the visual fixation point subset. Divide the total duration by the number of data in the duration subset, that is, use the total duration as the dividend and the number of data in the duration subset as the divisor. The result of the division is the unit duration intensity of the attention distribution unit.

[0038] The cognitive contribution of a unit is calculated by fusing the unit distribution density and unit persistence intensity. Then, the cognitive distribution feature value is calculated by integrating all unit cognitive contributions. Specifically, this involves fusing the unit distribution density and unit persistence intensity of each attentional distribution unit to obtain its cognitive contribution. The specific calculation method is as follows: the unit distribution density is multiplied by a preset density weight, and then the unit persistence intensity is multiplied by a preset intensity weight. That is, the unit distribution density is first used as the multiplicand, and the preset density weight is used as the multiplier; the two are multiplied to obtain the density contribution. Then, the unit persistence intensity is used as the multiplicand, and the preset intensity weight is used as the multiplicand; the two are multiplied to obtain the intensity contribution. Finally, the density contribution and intensity contribution are added together to obtain the unit cognitive contribution. The density weight and intensity weight are preset fixed values ​​based on the attentional cognitive patterns of children aged 3 to 8 years, and their sum equals 1. Specifically, the density weight ranges from 0.5 to 0.7, and the intensity weight ranges from 0.3 to 0.5. For example, a density weight of 0.6 and an intensity weight of 0.4 are used to balance the influence of distribution density and sustained intensity on cognitive contribution, ensuring that cognitive contribution can comprehensively reflect the child's attentional state in the unit, thus solving the problem that existing technologies cannot comprehensively assess the child's attentional cognitive state. The cognitive contribution of all attentional distribution units is integrated to calculate the cognitive distribution characteristic value. Specifically, this involves using a successive summation algorithm to successively add the cognitive contribution of all attentional distribution units. First, the cognitive contribution of the first unit is added to the cognitive contribution of the second unit to obtain a temporary total contribution. Then, the temporary total contribution is added to the cognitive contribution of the third unit, and so on, until the cognitive contribution of all attentional distribution units has been traversed, ultimately obtaining the total cognitive contribution of all units. Finally, the total cognitive contribution is divided by the total number of attentional distribution units, that is, the total cognitive contribution is used as the dividend and the total number of attentional distribution units is used as the divisor. The result is the cognitive distribution characteristic value.

[0039] This embodiment generates attention distribution units by using the minimum bounding rectangle, and then obtains the unit distribution density, unit persistence intensity, and cognitive distribution characteristic values ​​through specific calculation methods and a clear weight value range. It can accurately capture the details of children's attention distribution when reading, identify the specific picture book area that children focus on, and quantify the density and duration of children's attention.

[0040] In a preferred embodiment of the present invention, the emotional state analysis model is optimized based on cognitive distribution feature values, and the optimized model is used to process the reading behavior data stream to obtain structured analysis results characterizing children's cognitive and emotional states, which may include:

[0041] The cognitive distribution feature values ​​are input into a pre-defined model parameter mapping function. The model parameter mapping function outputs the emotional feature threshold adjustment and attention weight adjustment of the emotional state analysis model. Specifically, this includes: firstly, constructing an emotional state analysis model. This model takes the visual fixation point sequence pattern features and physiological signal rhythm features from children's reading behavior data stream as input, and the cognitive attention quantification value, emotional category classification probability, and emotional intensity level index as output. The model adopts a deep learning network structure combining convolutional neural networks and recurrent neural networks, including an input layer, a feature extraction layer, a fully connected layer, and an output layer. The input layer receives the input feature data and performs standardization processing. The standardization processing method is input feature... The value is subtracted from the feature mean and then divided by the feature standard deviation. The preset feature mean range is 0.2 to 0.4, and the preset feature standard deviation range is 0.1 to 0.3. The feature extraction layer extracts spatial features of visual fixation sequence through convolutional layers and temporal features through recurrent layers. At the same time, it extracts key information of physiological signal rhythm features. The preset convolutional kernel size of the convolutional layer is 3×3, and the preset number of hidden neurons in the recurrent layer is 64. The fully connected layer is used to fuse the extracted features. The preset number of neurons in the fully connected layer is 32. The output layer is used to output the final analysis results. The preset activation function of the output layer is the sigmoid function to ensure that the model can accurately capture the correlation features between children's cognition and emotion.

[0042] The constructed emotional state analysis model was trained using a large dataset of reading samples from children aged 3 to 8. This dataset covers data from children of different ages and reading habits, specifically including visual fixation sequence data, physiological signal data, and corresponding manually labeled cognitive focus, emotion category, and emotion intensity data. The total number of samples was no less than 1000 sets, with each set containing no less than 500 fixation points and corresponding physiological signal data. During training, the visual fixation sequence data and physiological signal data from the sample data were input into the model. The model's forward propagation was used to calculate the predicted output. The predicted output was then compared with the manually labeled true results, and the mean squared error algorithm was used to calculate the error between the two. The calculation method is to square the difference between the predicted output and the actual result, then sum all the squared differences and divide by the number of sample data. The gradient descent algorithm is used to iteratively adjust the parameters of the model, with an iteration step size of 0.001 and a preset maximum number of iterations of 1000. The prediction error is continuously reduced until the model's prediction accuracy reaches a preset training termination condition of more than 90%, at which point training stops and the initial sentiment state analysis model is obtained. The initial model stores the original sentiment feature threshold and the original attention weight matrix. The original sentiment feature threshold ranges from 0.3 to 0.5, with a specific preset value of 0.4. The original attention weight matrix is ​​a 2×2 matrix with matrix element values ​​ranging from 0.2 to 0.8, with a specific preset matrix of [0.5, 0.3] and [0.4, 0.6], which is used for subsequent model optimization.

[0043] The calculated cognitive distribution feature values ​​are input into a pre-defined model parameter mapping function. This function is pre-defined based on the correspondence between the cognitive characteristics of children aged 3 to 8 and model parameters; specifically, it is a linear mapping function with the following expression: ,in This represents the amount of model parameter adjustment. Here, k represents the cognitive distribution feature value, 'a' represents the preset proportional coefficient, and 'b' represents the preset baseline offset. This function accurately converts the cognitive distribution feature value into corresponding model parameter adjustments. Specifically, after receiving the cognitive distribution feature value, the model parameter mapping function calculates it using a preset linear mapping rule. This rule consists of two independent linear mapping relationships, corresponding to the emotional feature threshold adjustment and the attention weight adjustment, respectively. The specific implementation process is as follows: The first linear mapping is used to calculate the emotional feature threshold adjustment, which is calculated as follows: the emotional feature threshold adjustment equals the cognitive distribution feature value multiplied by the preset coefficient 'a', plus the preset offset 'b', where the preset coefficient 'a' ranges from 0 to 1. The first linear mapping is used to calculate the attention weight adjustment. The attention weight adjustment is a 2×2 matrix. The calculation method for each matrix element is that the adjustment corresponding to the element is equal to the cognitive distribution feature value multiplied by the preset coefficient c, plus the preset offset d. The preset coefficient c ranges from 0.05 to 0.15, and is specifically preset to 0.1. The preset offset d ranges from -0.03 to 0.03, and is specifically preset to 0. Through this linear mapping rule, the emotional feature threshold adjustment and attention weight adjustment of the emotional state analysis model are output.

[0044] Based on the emotional feature threshold adjustment, the preset original emotional feature thresholds in the emotional state analysis model are read; the original emotional feature thresholds are added to the emotional feature threshold adjustment to generate updated emotional feature thresholds; based on the attention weight adjustment, the preset original attention weight matrix in the emotional state analysis model is read, specifically including: based on the emotional feature threshold adjustment output by the model parameter mapping function, the preset original emotional feature thresholds in the model are read through the preset parameter reading channel inside the emotional state analysis model. This parameter reading channel is a dedicated data transmission channel for model integration, originating from the same source as the parameter writing channel during model training, ensuring the stability and accuracy of parameter reading. This original emotional feature threshold is the model training... After training, a fixed parameter is preset and stored, specifically a value of 0.4. This serves as the baseline threshold for filtering sentiment features during the initial model runtime. Parameter validation is performed concurrently during the reading process. Specifically, the original sentiment feature threshold is compared in real-time with a preset range of 0.3 to 0.5. If the read value is less than 0.3 or greater than 0.5, a parameter rereading instruction is immediately triggered. If the value still does not meet the range requirement after three rereads, the parameter update is terminated, and a parameter error is indicated to prevent reading errors or parameter anomalies from affecting subsequent update processes. The original sentiment feature threshold is then added to the sentiment feature threshold adjustment amount. Specifically, this calculation involves first retrieving the original sentiment feature threshold value from the model's internal processing unit. The emotional feature threshold is first determined, and then the adjustment amount of the emotional feature threshold output by the model parameter mapping function is retrieved. This ensures that both parameters are loaded into the computation unit simultaneously. The original emotional feature threshold is used as the addend, and the adjusted emotional feature threshold is used as the addend. The two are directly added together, and immediately after the calculation, a range check is performed. Specifically, the calculated value is compared with a reasonable range of 0.3 to 0.5. If the value is greater than 0.5, it is automatically corrected to 0.5; if it is less than 0.3, it is automatically corrected to 0.3; if it is within the range, the result is retained. This yields the updated emotional feature threshold, ensuring that the model can operate normally after the parameter update and avoiding bias in emotional feature selection due to parameters exceeding the reasonable range. This is based on the model parameter mapping... The attention weight adjustment output by the function is read from the preset original attention weight matrix in the sentiment state analysis model through the same parameter reading channel. This original attention weight matrix is ​​a 2×2 matrix that is preset and stored after the model training is completed. Specifically, the preset matrix is ​​[0.5, 0.3], [0.4, 0.6], which is used to balance the influence of visual gaze point sequence pattern features and physiological signal rhythm features on the analysis results. During the reading process, dual verification is performed simultaneously. The first verification is matrix dimension verification, which determines whether the read matrix is ​​2×2 in dimension. If the dimension does not match, a matrix reconstruction instruction is triggered to reread the original matrix. The second verification is element value verification, which compares each of the four elements in the matrix one by one to confirm that each element is between 0.2 and 0.If any single element outside the range of 8 is found, the entire original attention weight matrix is ​​reread.

[0045] Each element in the original attention weight matrix is ​​added to the corresponding adjustment value in the attention weight adjustment matrix to generate an updated attention weight matrix. Specifically, this involves clarifying the positional correspondence between the original attention weight matrix and the attention weight adjustment. The element in the first row and first column of the original attention weight matrix corresponds to the element in the first row and first column of the attention weight adjustment; the element in the first row and second column of the original attention weight matrix corresponds to the element in the first row and second column of the attention weight adjustment; the element in the second row and first column of the original attention weight matrix corresponds to the element in the second row and first column of the attention weight adjustment; and the element in the second row and second column of the original attention weight matrix corresponds to the element in the second row and second column of the attention weight adjustment. This correspondence is pre-stored in the model's computation unit to ensure that the positional correspondence is accurate. Then, the model's computational unit performs element-wise addition. Specifically, the first row and first column element of the original attention weight matrix is ​​added to the first row and first column element of the attention weight adjustment, resulting in the updated first row and first column element of the attention weight matrix. For example, if the original element is 0.5 and the corresponding adjustment is 0.02, the updated element is 0.52. The second row and second column element of the original attention weight matrix is ​​added to the first row and second column element of the attention weight adjustment, resulting in the updated first row and second column element of the attention weight matrix. For example, if the original element is 0.3 and the corresponding adjustment is 0.01, the updated element is 0.31. The element in the second row and first column of the attention weight matrix is ​​added to the element in the second row and first column of the attention weight adjustment to obtain the element in the second row and first column of the updated attention weight matrix. For example, if the original element is 0.4 and the corresponding adjustment is 0.03, the element after the operation is 0.43. The element in the second row and second column of the original attention weight matrix is ​​added to the element in the second row and second column of the attention weight adjustment to obtain the element in the second row and second column of the updated attention weight matrix. For example, if the original element is 0.6 and the corresponding adjustment is 0.02, the element after the operation is 0.62. After all four element operations are completed, the operation unit integrates them to form the complete updated attention weight matrix.

[0046] The updated sentiment feature thresholds and the updated attention weight matrix are written to the corresponding storage locations in the sentiment state analysis model to complete the model parameter update and obtain the optimized sentiment state analysis model. Specifically, this includes: determining the corresponding storage locations of the parameters in the sentiment state analysis model; the original sentiment feature thresholds are stored in the threshold storage area of ​​the model parameter storage unit, which is an independent partition using read-only locking, unlocked only during parameter updates; the original sentiment feature thresholds range from 0.3 to 0.5; the original attention weight matrix is ​​stored in the weight matrix storage area of ​​the model parameter storage unit, which is also an independent partition, stored separately from the threshold storage area to avoid parameter confusion; the updated sentiment feature thresholds are written to the threshold storage area through the model's parameter writing channel; the updated sentiment feature thresholds range from 0.3 to 0.5, and during writing, the threshold storage area is unlocked first, and an overwrite operation is performed to completely overwrite the original sentiment values ​​in that storage area. After the feature threshold is written, the threshold storage area is immediately locked, and a write verification is performed. Specifically, the verification operation involves rereading the parameters in the threshold storage area through the parameter reading channel and comparing the read parameters with the updated sentiment feature threshold to ensure that the two are completely consistent. If they are inconsistent, the storage area is unlocked and rewritten. If they are still inconsistent after three repetitions, the update is terminated and a write error is indicated. At the same time, the updated attention weight matrix is ​​written to the weight matrix storage area through the same parameter writing channel. The writing process is the same as that for sentiment feature threshold writing: first, the storage area is unlocked, an overwrite write is performed, the storage area is locked, and then the matrix in the weight matrix storage area is reread through the parameter reading channel. The matrix dimension and element values ​​are compared element by element to ensure that they are completely consistent with the updated attention weight matrix to ensure that the parameter writing is accurate. After all parameter writing verifications pass, the model automatically triggers a parameter update confirmation command to complete the parameter update of the sentiment state analysis model and obtain the optimized sentiment state analysis model.

[0047] The reading behavior data stream is input into the optimized emotional state analysis model to extract visual gaze point sequence pattern features and physiological signal rhythm features contained in the reading behavior data stream. Specifically, this includes: inputting the collected raw reading behavior data stream into the optimized emotional state analysis model, and the model performing feature extraction processing on the reading behavior data stream. The extraction process adopts a hierarchical extraction method, with a preset hierarchical extraction threshold of feature intensity 0.2. Features below this threshold will be discarded. On the one hand, the visual gaze point sequence pattern features contained in the data stream are extracted, specifically including the movement trajectory of the gaze point, the switching frequency between different attention distribution units, the residence time distribution in each unit, and the movement speed of the gaze point. Features such as frequency of attention distribution are extracted. The switching frequency is calculated as the number of times the attention distribution unit is switched per unit time, and the movement speed is calculated as the distance between two adjacent fixation points divided by the time difference between the two fixation points. On the other hand, the physiological signal rhythm features contained in the data stream are extracted. The physiological signals come from the physiological acquisition devices on the smart reading terminal, specifically the heart rate sensor and the skin conductance sensor. The collected physiological signals include the child's heart rate, skin conductance signals, etc. Features such as rhythm changes, fluctuation amplitude, and peak frequency of these physiological signals are extracted. The fluctuation amplitude is calculated as the maximum value of the physiological signal minus the minimum value of the physiological signal, and the peak frequency is calculated as the number of times the peak value of the physiological signal occurs per unit time.

[0048] Based on visual gaze point sequence pattern features and physiological signal rhythm features, the optimized emotion state analysis model calculates the cognitive attention quantification value, emotion category assignment probability, and emotion intensity level index. Specifically, the optimized emotion state analysis model performs comprehensive analysis and calculation based on the extracted visual gaze point sequence pattern features and physiological signal rhythm features to obtain the corresponding cognitive attention quantification value, emotion category assignment probability, and emotion intensity level index. The cognitive attention quantification value is calculated by multiplying the feature value corresponding to the visual gaze point sequence pattern feature by the corresponding weight in the updated attention weight matrix, adding the feature value corresponding to the physiological signal rhythm feature, multiplying it by the corresponding weight in the updated attention weight matrix, and then filtering and adjusting it through an updated emotion feature threshold. Specifically, if the weighted sum is greater than or equal to the updated emotion feature threshold, the original weighted sum is retained; otherwise, it is set to 0, resulting in the final cognitive attention quantification value, which ranges from 0 to 1. The calculation of the emotion category assignment probability... The method involves fusing the extracted two types of features and then calculating the result using a preset softmax classification function. This function first calculates the feature output value for each emotion category, then uses that value as the exponent of an exponential function to calculate the exponent value. Finally, it divides the exponent value of each emotion category by the sum of all emotion category exponent values ​​to obtain the probability of belonging to that emotion category. The probability values ​​of children belonging to different emotion categories (such as pleasure, boredom, curiosity, and irritability) are calculated in this way, and the sum of all emotion category probability values ​​is 1. The emotion intensity level index is calculated by comprehensively quantifying the maximum probability of belonging to the emotion category, combined with the fluctuations in visual fixation point sequence pattern features and physiological signal rhythm features. Specifically, the maximum probability of belonging to the emotion category is multiplied by a feature fluctuation coefficient, which is the average of the physiological signal fluctuation amplitude and the visual fixation point movement speed, and then divided by a preset baseline coefficient of 0.5. The index ranges from 0 to 1.

[0049] The method combines cognitive attention metrics, emotion category probability, and emotion intensity level index to generate structured analysis results. Specifically, it integrates the calculated cognitive attention metrics, emotion category probability, and emotion intensity level index, and arranges them according to a preset structured format. This preset structured format arranges the fields in a fixed order. First, the cognitive attention metrics are presented, retaining three decimal places. Then, each emotion category and its corresponding probability are presented in descending order of probability value, with each emotion category corresponding to one line, the emotion category name first, followed by the probability, and retaining three decimal places. Finally, the emotion intensity level index is presented, retaining three decimal places. No additional formatting is required; the results are simply arranged in this fixed order to generate structured analysis results representing children's cognitive and emotional states.

[0050] This embodiment, by constructing an emotional state analysis model in detail, clarifies the model's network structure, the function of each layer, and the input and output parameters, and details the sample data requirements, training process, and termination conditions for model training, ensuring that the model can accurately analyze children's cognitive and emotional states.

[0051] like Figure 2 As shown, in another preferred embodiment of the present invention, based on the structured analysis results, matching and mapping are performed to obtain a set of AI picture book strategy parameters, which may include:

[0052] Based on the probability of emotional category attribution, a preset emotional guidance strategy library is matched to obtain the most suitable basic emotional guidance strategy for the current emotional category. Specifically, this includes: calling the preset emotional guidance strategy library, which is constructed by first collecting effective guidance strategies corresponding to different emotional categories based on the emotional cognitive patterns of children aged 3 to 8 and the guidance needs of picture book reading through a large number of children's emotional cognitive experiments. Each emotional category corresponds to a set of exclusive basic emotional guidance strategies. Each set of basic emotional guidance strategies clearly includes core content such as visual presentation parameters, content rhythm parameters, and interaction method parameters. Among them, visual presentation parameters are used to guide the design of visual elements in picture books, content rhythm parameters are used to guide the presentation rhythm of picture book content, and interaction method parameters are used to guide the interaction design between picture books and children. After collection, all guidance strategies are classified and organized, and labeled according to emotional category. After labeling, multiple groups of children aged 3 to 8 are organized to conduct reading tests. By observing the children's state and feedback during reading, the suitability of each guidance strategy is verified. Strategies with poor suitability are adjusted and optimized. After multiple rounds of testing and optimization, a complete emotional guidance strategy library is constructed. After being stored in the library, it is classified and stored according to emotional category for easy matching and calling in the future. After the strategy library is invoked, the child's current dominant emotion category is determined based on the probability of emotion category attribution. The determination method is to select the emotion category with the highest probability value as the current dominant emotion category. If two or more emotion categories have the same probability value and both are the maximum value, then the preset emotion category with higher priority is selected as the dominant emotion category. The preset emotion priority order is curiosity, pleasure, boredom, and irritability. This priority order is preset based on the child's reading experience and the effect of emotional guidance, which can prioritize the cultivation of the child's reading interest. After determining the current dominant emotion category, a one-to-one precise matching is performed from the preset emotion guidance strategy library based on the dominant emotion category. During the matching, the matching algorithm built into the strategy library compares the dominant emotion category with the guidance strategies corresponding to various emotions in the library to obtain the basic emotion guidance strategy that is most suitable for the current dominant emotion category.

[0053] Based on the cognitive attention metric, the visual presentation parameters in the basic emotional guidance strategy are adjusted to generate a primary strategy parameter set. Based on the emotional intensity level index, the intensity of the primary strategy parameter set is modulated to obtain a complete strategy parameter set. Specifically, this includes: obtaining the visual presentation parameters in the basic emotional guidance strategy, which include the baseline values ​​for visual element complexity, color saturation, dynamic effects, and text-image ratio. These baseline values ​​are preset based on the general needs of the corresponding emotional categories and can meet the basic emotional guidance needs, but they cannot adapt to the child's current attention level. Therefore, they need to be adjusted specifically based on the cognitive attention metric. The adjustment rules are pre-set based on the cognitive attention patterns of children aged 3 to 8, enabling precise matching of visual presentation parameters with children's attention levels. The specific adjustment methods are as follows: the visual element complexity baseline value is added to the cognitive attention quantification value multiplied by 0.2 to obtain the adjusted visual element complexity parameter; the color saturation baseline value is added to the cognitive attention quantification value multiplied by 0.1 to obtain the adjusted color saturation parameter; the dynamic effect baseline value is added to the cognitive attention quantification value multiplied by 0.1 to obtain the adjusted dynamic effect parameter; and the image-text ratio baseline value is added to the cognitive attention quantification value multiplied by 0.15 to obtain the adjusted image-text ratio parameter. After all visual presentation parameters are adjusted, they are integrated with other core parameters in the basic emotional guidance strategy, including content rhythm parameters and interaction method parameters. During integration, parameters are categorized according to their functional attributes to ensure an orderly and clear arrangement, generating a primary strategy parameter set. This primary strategy parameter set can adapt to the child's current attention level, but it does not consider the influence of the child's current emotional intensity and requires further adjustment.

[0054] Based on the emotional intensity level index, the primary strategy parameter set is modulated. The modulation rules are pre-defined to achieve precise matching between parameters and emotional intensity. Specifically, each parameter in the primary strategy parameter set is multiplied by the emotional intensity level index and then added to a corresponding baseline offset. The baseline offset is a fixed value preset according to the parameter type: 0.02 for visual element complexity, 0.01 for color saturation, 0.01 for dynamic effects, 0.02 for text-image ratio, and 0 for content rhythm and interaction methods. This ensures that the modulation effect of different parameter types aligns with the emotional guidance requirements. During the modulation process, each parameter in the primary strategy parameter set is calculated and modulated one by one. After the calculation, the range of each modulated parameter is checked to ensure that each modulated parameter is within the preset reasonable range, avoiding deviations in subsequent visual content generation due to parameter anomalies. After all parameters are modulated, all modulated parameters are integrated, duplicate parameters are removed, and they are arranged according to the preset parameter classification order to obtain the complete strategy parameter set.

[0055] A complete set of strategy parameters is integrated to generate an AI picture book strategy parameter set. This includes: obtaining the complete set of strategy parameters and categorizing all parameters according to their functional attributes, specifically into four categories: visual element control parameters, narrative rhythm control parameters, emotional guidance control parameters, and interactive effect control parameters. Real-time verification is performed during the categorization process, checking the functional attributes of each parameter one by one to ensure accurate classification without errors or omissions. Specifically, visual element control parameters control the design of the picture book's visual elements, including adjusted visual element complexity parameters, color saturation parameters, and text-image ratio parameters. Narrative rhythm control parameters control the presentation rhythm of the picture book content, including content rhythm parameters and page transition speed. The system includes parameters such as degree of sensitivity; emotional guidance control parameters, which guide children's emotions, including emotional guidance strategy identifiers and interaction method parameters; and interaction effect control parameters, which control the interaction effect between the picture book and children, including dynamic effect parameters and feedback delay parameters. After classification, the parameters are arranged in a preset order: visual element control parameters, narrative rhythm control parameters, emotional guidance control parameters, and interaction effect control parameters. This order is based on the priority of the parameters in the picture book visual content generation process, which facilitates the subsequent parsing and use of parameters. During the arrangement process, each parameter is validated twice to verify whether its value is within the preset reasonable range and to verify the consistency of the parameter names to avoid confusion and ensure the standardization of parameters. Subsequently, all the arranged parameters are integrated in a preset format, which is set based on the parameter reading requirements of the subsequent picture book visual content generation process, enabling rapid reading and parsing of parameters. After integration, an AI picture book strategy parameter set is generated.

[0056] This embodiment solves the problem that picture book strategies are fixed and cannot adapt to children's real-time emotional and cognitive states by using structured analysis of core parameters and precise matching of emotional strategies, so that picture book guidance strategies can fit children's real reading state.

[0057] In a preferred embodiment of the present invention, the basic visual composition and frame layout of the picture book page are generated according to the AI ​​picture book strategy parameter set; based on the emotional guidance parameters in the AI ​​picture book strategy parameter set, the position, angle, and shape adjustment rules of each component of the basic visual composition are determined; and the basic visual composition and frame layout are moved, rotated, and adapted according to the shape adjustment rules to generate picture book visual content, which may include:

[0058] The algorithm parses visual element complexity adjustment coefficients, narrative rhythm control parameters, emotional guidance strategy identifiers, dynamic effect intensity parameters, and color saturation adjustment parameters from the AI ​​picture book strategy parameter set. Specifically, this includes: obtaining the AI ​​picture book strategy parameter set generated in Example 1, which contains all the core control parameters required for generating picture book visual content, ensuring the parameter set is complete and without omissions; employing a preset key-value pair matching parsing algorithm, the specific calculation process involves first initializing the algorithm, pre-setting target key names for five core parameters, corresponding to the visual element complexity adjustment coefficient, narrative rhythm control parameter, emotional guidance strategy identifier, dynamic effect intensity parameter, and color saturation adjustment parameter, respectively, and setting the key-value comparison threshold to 0.95. This threshold is used to determine the accuracy of key name matching, avoiding matching errors due to minor differences in key names; after initialization, the algorithm reads all key-value pair data from the AI ​​picture book strategy parameter set, and according to a preset byte splitting rule, splits each key-value pair into a key name string and a value (or identifier) ​​content. The splitting rule uses a preset separator as the splitting node. The default separator is a comma. After splitting, the key name strings are processed to remove spaces and ensure consistent capitalization, preventing format differences from affecting the comparison results. Following a preset parsing order, the processed key name strings are sequentially compared with the target key names of the five core parameters. The calculation method involves counting the number of identical characters in the two key name strings and then dividing that number by the total number of characters in the longer string to obtain the key name similarity value. The calculated similarity value is compared with a preset comparison threshold of 0.95. If the similarity value is greater than or equal to 0.95, the key name is considered a successful match, and the corresponding value (or identifier) ​​is extracted as the initial extraction value for that core parameter. If the similarity value is less than 0.95, the key name match is considered a failure, and a key name re-matching instruction is immediately triggered. The key name string is then compared again with the target key names of the five core parameters to eliminate matching failures caused by character order deviations. If the second comparison still fails, it is determined that the core parameter is missing from the AI ​​picture book strategy parameter set, triggering an exception prompt instruction and notifying staff to investigate the parameter.

[0059] The preset parsing order is: visual element complexity adjustment coefficient, narrative rhythm control parameter, emotional guidance strategy identifier, dynamic effect intensity parameter, and color saturation adjustment parameter, ensuring that the parsing process is orderly, complete, and without deviation. During parsing, the above algorithm is first used to read, split, compare, and extract initial values ​​of key-value pairs. Then, the extracted initial values ​​are converted to standard formats according to preset data format conversion rules. Specifically, the visual element complexity adjustment coefficient, dynamic effect intensity parameter, and color saturation adjustment parameter are converted to decimal format, retaining two decimal places. The narrative rhythm control parameter is converted to the standard character format of the corresponding level, and the emotional guidance strategy identifier is converted to the preset standard identifier format. After conversion, each data item is compared and confirmed with the five preset core parameter names one by one. If the comparison is successful, it is determined as the final target parameter. At the same time, each extracted parameter is checked for range or legality to ensure that the parameter is valid and usable, avoiding deviations in subsequent visual content generation due to parameter abnormalities. Among them, the visual element complexity adjustment coefficient is used to control the number and detail richness of visual elements on the picture book page. During verification, it is determined whether the coefficient is within the preset range. The parameter ranges from 0.3 to 0.9. If it exceeds this range, it is considered an abnormal parameter, and a re-parse or parameter correction instruction is immediately triggered. The narrative rhythm control parameter controls the content density and presentation rhythm of the picture book page. During verification, it is determined whether the parameter conforms to the three levels of soothing, moderate, and rapid and their corresponding value ranges to ensure that the parameter can effectively control the rhythm of the picture book content. The emotional guidance strategy indicator is used to determine the visual adjustment direction that matches the child's current emotion. During verification, it is determined whether the indicator is a preset legal indicator. Legal indicators include the joy indicator, boredom indicator, curiosity indicator, and irritability indicator. Each indicator corresponds to a clear emotional guidance direction to ensure that the indicator can accurately match the emotional guidance needs. The dynamic effect intensity parameter controls the prominence of the dynamic effect of visual elements. During verification, it is determined whether the parameter is within the preset value range of 0.1 to 0.7 to ensure that the dynamic effect is moderate and conforms to the child's visual cognitive habits. The color saturation adjustment parameter controls the color vibrancy of visual elements. During verification, it is determined whether the parameter is within the preset value range of 0.4 to 0.8 to ensure that the color effect is comfortable and matches the child's emotional state.

[0060] Based on the visual element complexity adjustment coefficient and narrative rhythm control parameters, a preset picture book material library is invoked to select matching character images, background scenes, and decorative objects, which are then combined to generate the basic visual composition of the picture book page. A preset page layout rule library is then invoked to select matching text and image arrangements and area proportions, generating the framework layout of the picture book page. Specifically, this includes: invoking the preset picture book material library, which is constructed by first collecting character images, background scenes, and decorative objects of different styles and levels of detail, tailored to the cognitive characteristics and reading preferences of children aged 3 to 8. Character images include animals, people, and cartoon characters; background scenes include family, nature, and school; and decorative objects include flowers, toys, and ornaments. The library also collects materials in various styles, such as cartoon, realistic, and minimalist, to meet different emotional guidance and narrative needs. After collection, all materials are categorized into three levels based on visual element complexity: simple, moderate, and rich; and into three categories based on narrative rhythm: soothing, moderate, and fast. Each material is labeled with its corresponding complexity level and suitable rhythm level for easy matching and selection later. After annotation, all materials are screened and optimized, removing those with blurry image quality or unsuitable styles. The material library is also regularly updated with new materials to ensure its richness and applicability. After multiple rounds of screening and optimization, a complete preset picture book material library is constructed. Once the material library is accessed, a coefficient is adjusted based on the complexity of the visual elements to determine the material's complexity level. A coefficient of 0.3 to 0.5 corresponds to a simple level of material, meaning simple character designs, simple background scenes, and a small number of decorative objects. A coefficient of 0.5 to 0.7 corresponds to a moderate level of material, meaning the character designs have basic details, the background scenes contain core elements, and the number of decorative objects is moderate. A coefficient of 0.7 to 0.9 corresponds to a rich level of material, meaning the character designs are rich in detail, the background scenes have complete elements, and the decorative objects are exquisite and diverse.

[0061] Based on the narrative rhythm control parameters, the appropriate rhythm for the materials is determined. A slow rhythm suits simple and gentle materials, a medium rhythm suits materials with a moderate style, and a fast rhythm suits simple and dynamic materials. Combining these requirements, matching character images, background scenes, and decorative objects are selected one by one from the picture book material library. During the selection process, a second verification is performed to check whether the selected materials meet the requirements for visual element complexity and whether they are suitable for the current narrative rhythm, ensuring that the selected materials accurately meet the requirements of both parameters. After selection, these materials are combined according to the picture book narrative logic. The character image is placed near the center of the image as the visual core, the background scene serves as the underlying background covering the entire basic area of ​​the page, and decorative objects are placed around the character image or in the blank areas of the page as embellishments. After combination, a rationality check is performed to examine whether the style of the materials is consistent, whether the proportions are harmonious, and whether there are any issues such as overlapping, occlusion, or excessive blank space. Any problems are adjusted and optimized to finally generate the basic visual composition of the picture book page.

[0062] The system utilizes a pre-defined page layout rule library. This library was built based on the visual cognitive patterns of children aged 3 to 8. Through extensive children's reading tests, it collected information on the presentation requirements of picture books corresponding to different levels of visual element complexity and narrative rhythm, and compiled various text and image arrangement methods and area proportion rules. The text and image arrangement methods include horizontally even arrangement, vertically even arrangement, arrangement with a focus on the core element and surrounding details, and simple and compact arrangement. The area proportion rules clearly define the proportion of character images, background scenes, decorative objects, and text areas. After organization, all rules were categorized and labeled according to the compatibility range of visual element complexity and narrative rhythm. Following labeling, multiple groups of children aged 3 to 8 underwent reading tests to observe their attention distribution and reading experience, verifying the rule suitability. Rules with poor suitability were adjusted and optimized. After multiple rounds of testing and optimization, a complete preset page layout rule library was constructed. After being added to the library, rules were categorized and stored according to parameter compatibility range for easy matching and retrieval later. When the rule library was accessed, the coefficients for adjusting visual element complexity and narrative rhythm control parameters were used to select matching layout rules from the library, clarifying the text and image arrangement and area proportions. The text and image arrangement was determined by the narrative rhythm control parameters, with a gentler rhythm using left-right... Alternatively, the elements can be arranged evenly vertically. For medium-paced scenes, a prominent core element with surrounding accents is used, while for fast-paced scenes, a simple and compact arrangement is employed. The area proportions are determined by the visual element complexity adjustment coefficient. When complexity is high, the character area occupies 30% to 40%, the background scene area 50% to 60%, and the decorative object area 5% to 10%. When complexity is low, the character area occupies 25% to 35%, the background scene area 45% to 55%, the decorative object area 10% to 15%, and the text area is fixed at 5% to 10%. This proportion range is based on the pre-set rules of children's visual cognition, ensuring that children can clearly read the text without affecting their viewing of visual elements. After determining the relevant rules, the framework layout of the picture book page is generated.

[0063] Based on the emotional guidance strategy identifiers and basic visual composition, the position offset reference values, angle deflection reference values, and shape scaling reference values ​​corresponding to each component are extracted from a preset emotional action mapping rule library. These are used as initial position adjustment rules, initial angle adjustment rules, and initial shape adjustment rules. Specifically, this includes calling the preset emotional action mapping rule library. The construction process of this rule library is as follows: Through a large number of emotional cognition experiments with children aged 3 to 8, children of different ages and personalities were invited to participate in the experiments. Real-time data on children's preferences for the position, angle, and shape of visual elements in picture books under different emotional states were collected. This was combined with children's... To address the emotional guidance needs, baseline values ​​for the positional offset, angle deflection, and shape scaling of visual elements corresponding to different emotional guidance strategy identifiers were determined. After establishing these baseline values, visual elements were categorized into three types: character images, background scenes, and decorative objects. Baseline values ​​for each type of visual element were compiled, and the corresponding emotional guidance strategy identifier and applicable range for each baseline value were labeled. Multiple rounds of experimental verification were conducted after labeling. The baseline values ​​were applied to the adjustment of visual elements in the picture book, and children's emotional feedback was observed. The baseline values ​​were adjusted and optimized to ensure they effectively guide children's emotions. After multiple rounds of verification and optimization, a complete set of preset emotional dynamics was constructed. A mapping rule base is created and stored according to emotion guidance strategy identifiers and visual element types for easy retrieval and retrieval later. After the rule base is accessed, corresponding baseline value groups are extracted based on the emotion guidance strategy identifiers. These baseline value groups contain position offset, angle deflection, and shape scaling baseline values ​​for different types of visual elements. For each component in the basic visual composition, including character image, background scene, and decorative objects, the corresponding position offset baseline value, angle deflection baseline value, and shape scaling baseline value are extracted one by one. The character image, as the visual core, has its baseline values ​​determined based on the emotion guidance direction. The "Joy" sign features a small positional offset, a near-zero angle shift, and a slight enlargement in shape, ensuring a bright and soft visual effect that aligns with joyful emotions. The "Curiosity" sign features a slightly larger positional offset, a 5-10 degree angle shift, and a moderately enlarged shape, ensuring a novel and varied visual effect that stimulates children's curiosity. The "Bored" sign features a medium positional offset, a 0-5 degree angle shift, and a slight reduction in shape, ensuring a simple and interesting visual effect that attracts children's attention. The "Annoyed" sign features the smallest positional offset, a 0-degree angle shift, and a slight reduction in shape, ensuring a soft and soothing visual effect that alleviates children's irritability.The baseline values ​​for the background scene are based on the character's image. The position offset baseline value is consistent with the character's image to ensure that the background and character positions are coordinated. The angle offset baseline value is fixed at 0 degrees to avoid background rotation affecting the visual experience. The shape scaling baseline value is adjusted synchronously according to the character's image scaling baseline value to ensure that the background and character proportions are coordinated and there is no stretching or shrinking abnormality. The baseline values ​​for decorative objects are adjusted according to the character's image and emotional guidance direction. The offset direction revolves around the character's image to ensure that the decorative objects play an embellishing role. The angle offset baseline value is 3 to 8 degrees, and the shape scaling baseline value is 0.8 to 1.1 to adapt to the corresponding emotional guidance needs and enhance the richness of the visual effect. After extraction, the three sets of baseline values ​​corresponding to each visual element are integrated to form the initial position adjustment rules, initial angle adjustment rules, and initial shape adjustment rules.

[0064] Based on the dynamic effect intensity parameter, proportional calculations are performed on the position offset reference value, angle deflection reference value, and shape scaling reference value to generate the final position movement, angle rotation, and shape adaptation amount for each component. Specifically, this includes: clarifying the role of the dynamic effect intensity parameter, which adjusts the magnitude of visual element adjustment; a larger parameter value results in a larger adjustment and a more pronounced dynamic effect, while a smaller parameter value results in a smaller adjustment and a smoother dynamic effect. This parameter enables precise and controllable dynamic effects, adapting to children's current emotional state and visual cognitive habits. For each visual element's position offset reference value, the final position movement is calculated by multiplying the position offset reference value by the dynamic effect intensity parameter, yielding the position movement amounts in both the horizontal and vertical axes. After calculation, a range check is performed to ensure the position movement is within the preset range of 0 to 0.05 page units, preventing excessive movement that could cause visual elements to exceed page boundaries. If the movement exceeds the acceptable range, the product of the dynamic effect intensity parameter and the baseline value should be adjusted appropriately to ensure that the movement is reasonable. For each visual element's angle deflection baseline value, the final angle rotation amount is calculated by multiplying the angle deflection baseline value of each visual element by the dynamic effect intensity parameter. After calculation, a range check is performed to control the rotation amount within the range of -15 degrees to 15 degrees, where a positive value represents clockwise rotation and a negative value represents counterclockwise rotation. This ensures that the visual element maintains a normal visual effect after rotation and does not appear inverted or excessively tilted. If the movement exceeds the acceptable range, it should be corrected to ensure that the rotation amount is reasonable. For each visual element's shape scaling baseline value, the final shape adaptation amount is calculated by multiplying the shape scaling baseline value of each visual element by the dynamic effect intensity parameter and then adding the preset baseline scaling value of 1.0 to the result. After calculation, a range check is performed to control the shape adaptation amount within the range of 0.8 to 1.3.

[0065] Based on the color saturation adjustment parameters, the original color values ​​of each element in the basic visual composition are superimposed to generate the adjusted color saturation value. Specifically, this includes: obtaining the original color value of each visual element in the basic visual composition; this value is the standard color saturation value inherent in the materials in the picture book resource library, with each visual element's original color value preset to around 0.5 to ensure the initial color is moderate and without abruptness, conforming to children's visual cognitive habits; superimposing the original color value of each visual element according to the obtained color saturation adjustment parameters; specifically, adding the color saturation adjustment parameter to the original color value of each visual element, and then subtracting the preset baseline correction value of 0.5 from the result; retaining two decimal places in the calculation process to avoid color deviation caused by calculation errors and ensure accurate color adjustment; and verifying the calculation result within a range to determine whether the adjusted color saturation value is within the preset range of 0.3 to 0.9. If the result is greater than 0.9, it is corrected to 0.9; if the result is less than 0.3, it is corrected to 0.3; if it is within a reasonable range, the result is retained to ensure that the color saturation is moderate, neither too bright and stimulating to children's eyes, nor too dim and affecting visual effects. After the color saturation of each visual element is adjusted, a comparison and verification is performed, comparing the adjusted color of each element with the colors of other elements one by one to ensure that the adjusted color not only meets the requirements of the color saturation adjustment parameters, but also coordinates and unifies with the colors of other visual elements, without any individual element colors being too abrupt, ensuring an overall visual effect that is comfortable, conforms to children's visual cognitive habits, and accurately matches children's current emotional state. After all the color saturation adjustments and verifications of all visual elements are completed, the adjusted color saturation value corresponding to each visual element is determined, clarifying the one-to-one correspondence between individual visual elements and their corresponding adjusted color saturation values, thus completing the generation of the adjusted color saturation value for each visual element.

[0066] The position shift, angle rotation, and shape adaptation are applied to the corresponding character images, background scenes, and decorative objects in the basic visual composition, respectively, to complete the translation, rotation, and scaling transformations in the coordinate system of the picture book page, resulting in the deformed and adapted visual composition elements. Specifically, this includes: clarifying the rules of the preset coordinate system of the picture book page. This coordinate system is pre-defined, with the upper left corner of the page as the origin, the horizontal axis pointing to the right as the positive direction, and the vertical axis pointing downwards as the positive direction. All transformations of visual elements are based on this coordinate system to ensure accurate transformation positions and avoid element offsets; for each visual element, deformation processing is performed sequentially in the order of translation, rotation, and scaling to ensure the transformation process is orderly and without deviation; during translation, the original coordinates of the element are added to the corresponding position shift to calculate the translated horizontal and vertical coordinates. After translation, it is verified whether the translated coordinates of the element are within its corresponding frame layout area. If they exceed the area, the position shift is adjusted appropriately to ensure that the element is completely within the preset area without offset or exceeding the boundary; during rotation, the coordinates of the element are adjusted to the corresponding position shift. The element's center coordinates serve as the rotation center point. The element is rotated by the corresponding angle, with the rotation direction determined by the sign of the angle: positive for clockwise and negative for counter-clockwise. After rotation, the element's boundary is checked to ensure it doesn't exceed the preset area, and there's no distortion or deformation to guarantee normal visual effects and a smooth reading experience for children. During scaling, the element's center coordinates are used as the scaling center point, and the element is scaled by the corresponding shape adaptation amount. The scaled width and height are calculated separately. After scaling, the element's size is checked to ensure it matches the frame layout area, preventing it from exceeding the area or creating excessive blank space. Details are also checked for clarity and no blurring or distortion to guarantee visual quality. After all three transformations are completed, each visual element is checked as a whole. The position, angle, and shape of each element after translation, rotation, and scaling are verified for accuracy and suitability, with no abnormal deviations. Elements with deviations are adjusted and optimized to obtain the final visually adapted elements.

[0067] The deformed and adapted visual elements are sequentially placed into the picture book page frame, and adjusted color saturation values ​​are assigned to generate the picture book's visual content. Specifically, this includes: determining the placement order of each deformed and adapted visual element according to the requirements of the picture book page frame layout; the placement order is background scene, character image, decorative objects, and text elements. This order is based on the hierarchy of visual presentation, ensuring that bottom-level elements are placed first and top-level elements are placed later to avoid element occlusion and ensure the rationality and integrity of the visual presentation; first, the deformed and adapted background scene elements are placed into the background scene area of ​​the frame layout, adjusting the size of the background scene elements to completely cover the background scene area, and simultaneously assigning adjusted color saturation values ​​to the background scene elements. After placement, verification is performed, checking each background scene for offset, stretching, and color uniformity, adjusting any problematic parts to ensure the background scene appears correctly; subsequently, the deformed and adapted character image elements are placed into the character image area of ​​the frame layout, adjusting the position of the character image to ensure the character's shape... The elephant is positioned in the center of the area, blending naturally with the background elements without any jarring effect. Simultaneously, adjusted color saturation values ​​are applied to the character image element. After placement, verification is performed to ensure the character image is unobstructed, unoffset, and that its shape and angle meet requirements, guaranteeing clear visibility and a good visual effect. Deformed decorative elements are then placed one by one into the decorative object area of ​​the frame layout according to preset positions. Each decorative object is assigned an adjusted color saturation value, and its position is adjusted during placement to ensure harmony and unity with the character image and background scene, with no overlap, obstruction, or positional offset. The quantity and layout meet the requirements of visual element complexity, effectively serving as embellishment. Based on narrative rhythm control parameters, corresponding text content is placed into the text area according to a preset text-image arrangement. The text content is preset according to the narrative logic of the picture book, aligning with the picture book theme and children's reading level. The text color saturation is adjusted synchronously according to color saturation adjustment parameters to ensure clear and legible text, harmony and unity with visual elements, and no impact on the overall visual effect. After all elements are placed, an overall visual verification is performed. Each element is checked to ensure that the page layout is reasonable, the element proportions are coordinated, the colors are consistent, there are no abnormal deviations, and the dynamic effects are appropriate. Once the verification is passed, the complete picture book visual content is generated.

[0068] This embodiment, by adjusting visual parameters in layers and accurately generating visual content for picture books, overcomes the shortcomings of picture books having a single visual form and being unable to change dynamically, so that the picture book's images, rhythm, and colors are all adapted to children's attention and emotional needs.

[0069] In a preferred embodiment of the present invention, during the presentation of the picture book's visual content, the reader's dynamic visual attention data stream is collected in real time and converted into updated cognitive distribution feature values, which are then fed back to the emotional state analysis model to dynamically adjust the generation of subsequent picture book visual content and achieve emotional interaction guidance. This may include:

[0070] This system collects dynamic visual attention data streams from readers during the presentation of picture book content in real time, generating a real-time gaze point sequence. Based on the coordinates of all real-time gaze points on the current picture book page, a real-time attention distribution plane is constructed. Specifically, this involves: first, activating the data acquisition device on the smart reading terminal, which consists of a high-definition camera and an eye-tracking device. The high-definition camera captures images of the child's face and eyes, while the eye-tracking device accurately captures the child's eye movement trajectory, ensuring the acquisition device is in real-time operation. After activation, initial calibration is performed, and the acquisition frequency is set to 100Hz. This frequency ensures accurate and real-time capture of dynamic visual attention changes during the child's viewing of picture book content, without data omissions or delays. The system then collects dynamic visual attention data streams from children during the presentation of picture book content, including core data such as the coordinates of real-time gaze points, gaze direction, and gaze duration. The coordinates of the real-time gaze points reflect the child's current focus position, the gaze direction reflects the child's gaze angle, and the gaze duration reflects the child's level of attention to a particular location. During the data collection process, the built-in algorithm of the acquisition device performs preliminary processing on the raw data to remove fuzzy and invalid data, ensuring the accuracy and effectiveness of the data. Then, the real-time acquired data is integrated in real-time according to timestamp order, with each timestamp corresponding to a complete set of real-time fixation point data. This integration generates a real-time fixation point sequence, which reflects the movement trajectory and focusing changes of children's attention in real time, providing basic data for subsequent attention analysis. Based on the position coordinates of all real-time fixation points in the real-time fixation point sequence on the current picture book page, a real-time attention distribution plane is constructed. First, the rules of the preset coordinate system of the picture book page are clarified. This coordinate system has the upper left corner of the page as the origin, and the values ​​of the horizontal and vertical coordinates are both in the range of 0 to 1, ensuring coordinate uniformity. Each real-time fixation point in the real-time fixation point sequence is mapped one by one onto the two-dimensional plane of the picture book page according to its corresponding position coordinates. As real-time fixation points are continuously acquired, the fixation point distribution on the two-dimensional plane is continuously updated, new fixation point coordinates are constantly added, and outdated and invalid fixation point data is deleted, constructing a real-time attention distribution plane that can reflect the current distribution of children's attention in real time.

[0071] Analyzing all real-time fixation points on the real-time attention distribution plane yields multiple real-time fixation point groups. The area covered by each real-time fixation point group is defined as a real-time attention aggregation region, generating a set of real-time attention aggregation regions. Specifically, this involves: employing a preset spatial aggregation judgment rule, which is pre-set based on the characteristics of children's attention distribution and can accurately identify attention aggregation regions. The specific rule is that if the straight-line distance between two real-time fixation points does not exceed 0.05 times the side length of a picture book page, and three or more consecutive real-time fixation points meet this distance condition, then these real-time fixation points are determined to form an aggregation state. This rule can avoid misjudgment due to accidental aggregation of a single or two fixation points; performing real-time spatial distribution analysis on all real-time fixation points on the real-time attention distribution plane, using a built-in spatial analysis algorithm to calculate the distance between adjacent real-time fixation points one by one, calculating the straight-line distance between two fixation points, and then comparing the calculation result with the preset 0.05 times the side length of a picture book page to filter out real-time fixation points that meet the aggregation state condition.

[0072] These real-time gaze points that meet the criteria are grouped according to their aggregation state. Grouping follows the principle that each gaze point belongs to only one group. An algorithm groups adjacent gaze points that meet the aggregation criteria together, forming a real-time gaze point group. During the grouping process, real-time verification is performed, checking whether the number of real-time gaze points in each group is not less than 3. If a group has fewer than 3 gaze points, it is considered an accidental aggregation, split, and regrouped to avoid misjudgment. After grouping, the entire spatial range covered by each real-time gaze point group is determined as a real-time attention aggregation region. This is determined by taking the maximum spatial range covered by the coordinates of all real-time gaze points in each group, specifically selecting the minimum and maximum values ​​of the horizontal and vertical coordinates. The spatial range determined by these four values ​​is the real-time attention aggregation region corresponding to that group. All determined real-time attention aggregation regions are then integrated and numbered chronologically to generate a set of real-time attention aggregation regions.

[0073] For each real-time attention-focused region, the minimum bounding rectangle of all real-time gaze points within the complete coverage area is calculated. Each minimum bounding rectangle is determined as a real-time attention distribution unit, generating a set of real-time attention distribution units. Specifically, for each real-time attention-focused region in the set, a coordinate traversal filtering algorithm within the region is used. This algorithm can accurately filter out the coordinates of all gaze points belonging to the region. The preset calculation threshold is 0.001 units outward from the boundary of the attention-focused region. This threshold ensures that all filtered gaze points belong to the region, without missing any valid gaze points. The specific calculation process is as follows: First, based on the spatial range of the current real-time attention-focused region, a real-time filtering coordinate interval is set. The horizontal coordinate interval is from the minimum horizontal coordinate of the region minus 0.001 to the maximum horizontal coordinate plus 0.001, and the vertical coordinate interval is from the minimum vertical coordinate of the region minus 0.001 to the maximum vertical coordinate plus 0.001. Then, the algorithm traverses all real-time gaze point coordinates in the current real-time gaze point sequence, filtering out the real-time gaze points whose coordinates fall within the real-time filtering interval, which are used as the core gaze point coordinates constituting the real-time attention-focused region.

[0074] During the extraction process, real-time secondary verification is performed. The distance between each selected real-time gaze point and the center of the real-time attention focus area is calculated. The calculation result is compared with a preset 0.05 times the side length of the picture book page. If the distance is no more than 0.05 times the side length of the picture book page, the verification is passed. If the distance of a gaze point exceeds this range, it is determined that the gaze point does not belong to the focus area and is removed to ensure that the selected gaze points are accurate and effective. Based on the coordinates of all extracted real-time gaze points, the minimum and maximum values ​​of the horizontal coordinate and the vertical coordinate are found. The specific search method is to first extract the coordinates of the real-time gaze points... All real-time gaze coordinates are arranged in timestamp order. The x-coordinate of the first real-time gaze point is selected as the initial minimum and maximum x-coordinate values. Then, the x-coordinates of each remaining real-time gaze point are traversed one by one. If the x-coordinate of the current gaze point is less than the initial minimum x-coordinate value, the initial minimum x-coordinate value is updated. If the x-coordinate of the current gaze point is greater than the initial maximum x-coordinate value, the initial maximum x-coordinate value is updated. After traversal, the minimum and maximum x-coordinate values ​​of all real-time gaze points in the region are obtained. The minimum and maximum y-coordinate values ​​are obtained in the same way to ensure accurate coordinate values.

[0075] Based on the selected minimum x-coordinate and maximum y-coordinate, the coordinates of the top-left vertex of the minimum bounding rectangle are determined, i.e., the top-left vertex coordinates are the minimum x-coordinate and maximum y-coordinate. Based on the maximum x-coordinate and minimum y-coordinate, the coordinates of the bottom-right vertex of the minimum bounding rectangle are determined, i.e., the bottom-right vertex coordinates are the maximum x-coordinate and minimum y-coordinate. Based on these two vertex coordinates, a minimum bounding rectangle that completely surrounds all real-time gaze points within the region and whose edges are parallel to the coordinate axes of the picture book page is generated. After generation, real-time deviation verification is performed. The preset parallel deviation threshold for the rectangle edge is 0.0001 units. The parallel deviation between the rectangle edge and the coordinate axes is calculated to ensure that the deviation does not exceed this threshold. At the same time, it is verified that all real-time gaze points are inside the rectangle, with no omissions or excesses. Each minimum bounding rectangle is directly determined as a real-time attention distribution unit. A real-time attention gathering region corresponds to a real-time attention distribution unit. All generated real-time attention distribution units are integrated one by one to generate a set of real-time attention distribution units.

[0076] Based on the real-time attention distribution unit set and the real-time gaze point sequence, the number of real-time gaze points and the gaze duration within each real-time attention distribution unit are counted. The real-time unit distribution density and real-time unit persistence intensity are calculated, generating a real-time unit distribution density set and a real-time unit persistence intensity set. Specifically, for each real-time attention distribution unit in the set, its specific boundary range is determined based on the coordinates of its top-left and bottom-right vertices. The horizontal boundary is from the horizontal coordinate of the top-left vertex to the horizontal coordinate of the bottom-right vertex, and the vertical boundary is from the vertical coordinate of the bottom-right vertex to the vertical coordinate of the top-left vertex. All real-time gaze points meeting the boundary range are selected and integrated in timestamp order to generate a real-time gaze point subset. This subset only contains gaze point data belonging to the current distribution unit. From the real-time acquired dynamic visual attention data stream, based on the timestamp of each real-time gaze point in this subset, the corresponding gaze duration is retrieved and integrated in timestamp order to generate a real-time persistence duration subset. After integration, real-time verification is performed, checking each gaze duration individually. The timestamps and corresponding durations of each fixation point are matched to ensure no missing data or errors, guaranteeing data accuracy. Based on a subset of real-time fixations, a point-by-point matching counting algorithm is used, with the counting order preset to timestamp order and a preset allowable deviation of 0 for secondary verification. The algorithm counts fixations one by one in the subset, calculating the total number of real-time fixations in the subset. After the count is completed, a real-time secondary verification is performed, recounting to ensure accurate results without over- or under-counting. Next, the area of ​​the real-time attention distribution unit is calculated. The side length of the picture book page is preset to 1 unit. The area is calculated by subtracting the minimum horizontal coordinate value from the maximum horizontal coordinate value of the unit, obtaining the horizontal coordinate difference. Then, the maximum vertical coordinate value of the unit is subtracted from the minimum vertical coordinate value, obtaining the vertical coordinate difference. Multiplying the horizontal and vertical coordinate differences gives the area of ​​the real-time attention distribution unit. The total number of real-time fixations is divided by the area of ​​the real-time attention distribution unit to obtain the real-time unit distribution density, which reflects the density of attention within the unit. Using the same method, the real-time unit distribution density is calculated for each real-time attention distribution unit. After the calculation is completed, all density values ​​are integrated in the order of distribution unit number to generate a real-time unit distribution density set.

[0077] Based on a subset of real-time durations, a value-by-value summation algorithm is used, with a preset summation precision of 0.1ms. The algorithm sums all duration values ​​in the subset one by one to obtain the total real-time duration, maintaining a 0.1ms precision during the summation process to avoid computational errors. Next, a point-by-point matching counting algorithm is used to count the number of data points in the subset, i.e., the total number of durations. After the count is completed, a reverse secondary check is performed, comparing the number of data points with the total number of fixation points in the real-time fixation point subset to ensure consistency and avoid data deviation. The total real-time duration is divided by the number of data points in the subset to obtain the real-time unit persistence intensity of that real-time attention distribution unit. This intensity reflects the child's sustained attention to that unit area. Using the same method, the real-time unit persistence intensity is calculated for each real-time attention distribution unit. After calculation, all intensity values ​​are integrated according to the distribution unit number order to generate a set of real-time unit persistence intensities.

[0078] The real-time unit cognitive contribution is calculated by weighting and fusing the corresponding values ​​belonging to the same real-time attention distribution unit from the real-time unit distribution density set and the real-time unit persistence intensity set. This generates a real-time unit cognitive contribution set. Specifically, for each real-time attention distribution unit, the real-time unit distribution density corresponding to that unit is extracted from the real-time unit distribution density set, and the real-time unit persistence intensity corresponding to that unit is extracted from the real-time unit persistence intensity set. During extraction, matching is performed using the unit number to ensure that the two values ​​correspond to the same real-time attention distribution unit, avoiding calculation errors caused by mismatches. A weighted fusion calculation method is used, which is pre-set based on the attentional cognitive patterns of children aged 3 to 8 years, enabling precise fusion of density and intensity. The specific calculation method is as follows. The density of the real-time unit distribution is multiplied by a preset density weight to obtain a density-weighted value. Then, the sustained intensity of the real-time unit is multiplied by a preset intensity weight to obtain an intensity-weighted value. The density-weighted value and the intensity-weighted value are added together to obtain the real-time cognitive contribution of the real-time attention distribution unit. The density weight and intensity weight are preset fixed values, and their sum equals 1. The density weight is 0.6 and the intensity weight is 0.4. This weight allocation is determined based on a large number of children's attention experiments and can more accurately reflect children's cognitive focus. Using the same method, the corresponding real-time unit cognitive contribution is calculated for each real-time attention distribution unit. After the calculation is completed, the units are integrated according to their numbering order to generate a set of real-time unit cognitive contributions.

[0079] The cognitive contributions of real-time units are summed, and the summation result is normalized to obtain updated cognitive distribution feature values. These updated cognitive distribution feature values ​​are then input into the emotional state analysis model, and the parameters of the emotional feature thresholds and attention weight matrices stored within the model are updated to obtain an emotional state analysis model optimized by real-time cognitive feedback. Specifically, this involves: sequentially summing the cognitive contributions of each real-time attention distribution unit to obtain the total cognitive contribution of all real-time units; normalizing this total cognitive contribution by dividing it by the total number of real-time attention distribution units to obtain updated cognitive distribution feature values; and finally, inputting these updated cognitive distribution feature values ​​into the emotional state analysis model. The emotional state analysis model updates the parameters of the emotional feature thresholds and attention weight matrices stored within the model. First, a detailed emotional state analysis model needs to be constructed. This model takes the visual gaze point sequence pattern features and physiological signal rhythm features from the children's reading behavior data stream as input, and the cognitive attention quantification, emotional category assignment probability, and emotional intensity level index as output. The model adopts a deep learning network structure combining convolutional neural networks and recurrent neural networks, including an input layer, a feature extraction layer, a fully connected layer, and an output layer. The input layer receives the input feature data and performs standardization processing. The standardization method is to subtract the feature mean from the input feature value and then divide by the feature standard deviation. The feature mean ranges from 0.2 to 0.4, and the feature standard deviation ranges from 0.1 to 0.3. The feature extraction layer extracts spatial features of the visual gaze sequence through convolutional layers, with the kernel size of the convolutional layers set to 3x3. It extracts time-series features through recurrent layers, with the number of hidden neurons in the recurrent layers set to 64, while also extracting key information of physiological signal rhythm features. The fully connected layer is used to fuse the extracted features, with the number of neurons in the fully connected layer set to 32. The output layer is used to output the final analysis results, and the activation function of the output layer is set to the sigmoid function to ensure that the model can accurately capture the correlation features between children's cognition and emotion.

[0080] The constructed emotional state analysis model was trained using a large amount of sample data from children aged 3 to 8 reading. The sample data covered relevant data from children of different ages and reading habits, including visual fixation sequence data, physiological signal data, and corresponding manually labeled cognitive attention, emotion category, and emotion intensity data. The total number of sample data was no less than 1,000 sets, and each set of sample data contained no less than 500 fixation points and corresponding physiological signal data. During training, visual gaze sequence data and physiological signal data from the sample data are input into the model. The model's forward propagation is used to calculate the predicted output. The predicted output is then compared with the manually labeled real results. The mean squared error algorithm is used to calculate the error between the two. The mean squared error is calculated as the square of the difference between the predicted output and the real result. All squared differences are then summed and divided by the number of sample data. The gradient descent algorithm is used to iteratively adjust the model parameters. The iteration step size is set to 0.001, and the maximum number of iterations is preset to 1000. The prediction error is continuously reduced until the model's prediction accuracy reaches the preset training termination condition, at which point training stops, and the initial sentiment analysis model is obtained. The initial model stores the original sentiment feature threshold and the original attention weight matrix. The original sentiment feature threshold ranges from 0.3 to 0.5, with a preset value of 0.4. The original attention weight matrix is ​​a 2x2 matrix with element values ​​ranging from 0.2 to 0.8. The preset matrix has the first element being 0.5 and the second element being 0.3, ending with the first row and the second row having the first element being 0.4 and the second element being 0.6, which are used for subsequent model optimization.

[0081] During parameter updates, the updated cognitive distribution feature values ​​are first input into a preset model parameter mapping function. This function is pre-defined based on the correspondence between the cognitive characteristics of children aged 3 to 8 and model parameters; specifically, it is a linear mapping function that accurately converts the cognitive distribution feature values ​​into corresponding model parameter adjustments. After receiving the cognitive distribution feature values, the model parameter mapping function performs calculations according to preset linear mapping rules. These rules consist of two independent linear mapping relationships, corresponding to the adjustment amounts for emotional feature thresholds and attention weights, respectively. The first linear mapping is used to calculate the emotional feature threshold adjustment. The calculation method is that the emotional feature threshold adjustment equals the cognitive distribution feature value multiplied by a preset coefficient 'a', plus a preset offset 'b'. The preset coefficient 'a' ranges from 0.1 to 0.3, specifically preset to 0.2, and the preset offset 'b' ranges from -0.05 to 0.05, specifically preset to 0. The second linear mapping is used to calculate the attention weight adjustment. The attention weight adjustment is a 2x2 matrix. The calculation method for each matrix element is that the adjustment corresponding to that element equals the cognitive distribution feature value multiplied by a preset coefficient 'c', plus a preset offset 'd'. The preset coefficient 'c' ranges from 0.05 to 0.15, specifically preset to 0.1, and the preset offset 'd' ranges from -0.03 to 0.03, specifically preset to 0. Through this linear mapping rule, the emotional feature threshold adjustment and attention weight adjustment of the emotional state analysis model are output.

[0082] Based on the emotional feature threshold adjustment, the pre-stored original emotional feature thresholds in the emotional state analysis model are read through a pre-defined parameter reading channel within the model. This parameter reading channel is a dedicated data transmission channel integrated into the model, originating from the same parameter writing channel used during model training, ensuring the stability and accuracy of parameter reading. The read original emotional feature thresholds are added to the emotional feature threshold adjustment to obtain the updated emotional feature thresholds. After the calculation, a range check is performed, controlling the updated emotional feature thresholds within a reasonable range of 0.3 to 0.5. If the value is greater than 0.5, it is automatically corrected to 0.5; if it is less than 0.3, it is automatically corrected to 0.3; if it is within the range, the calculation result is retained. Based on the attention weight adjustment, the pre-defined original attention weight matrix in the emotional state analysis model is read through the same parameter reading channel, clarifying the positional correspondence between the original attention weight matrix and the attention weight adjustment. The first row of the original attention weight matrix... The first column of elements corresponds to the first row and first column of the attention weight adjustment amount. The first row and second column of the original attention weight matrix corresponds to the first row and second column of the attention weight adjustment amount. The second row and first column of the original attention weight matrix corresponds to the second row and first column of the attention weight adjustment amount. The second row and second column of the original attention weight matrix corresponds to the second row and second column of the attention weight adjustment amount. The correspondence is pre-stored in the model's computation unit to ensure that the positional correspondence is without deviation. Then, each element in the original attention weight matrix is ​​added to the corresponding adjustment value in the attention weight adjustment amount to generate the updated attention weight matrix.

[0083] The updated sentiment feature thresholds and the updated attention weight matrix are written to the corresponding storage locations in the sentiment state analysis model. The original sentiment feature thresholds in the sentiment state analysis model are stored in the threshold storage area of ​​the model parameter storage unit. This storage area is an independent partition and uses a read-only locking method, which is only unlocked when the parameters are updated. The original attention weight matrix is ​​stored in the weight matrix storage area of ​​the model parameter storage unit. This storage area is also an independent partition and is stored separately from the threshold storage area to avoid parameter confusion. When writing, the corresponding storage area is unlocked first, and an overwrite write operation is performed to completely overwrite the original parameters. After the write is completed, the storage area is immediately locked, and a write verification is performed. The parameters in the storage area are reread through the parameter reading channel and compared with the updated parameters to ensure that the two are completely consistent. After all parameter write verifications pass, the parameter update of the sentiment state analysis model is completed, and the sentiment state analysis model optimized by real-time cognitive feedback is obtained.

[0084] The real-time gaze point sequence is input into an emotion state analysis model optimized by real-time cognitive feedback. Through processing, updated cognitive attention quantification, updated emotion category attribution probability, and updated emotion intensity level index are obtained. These are combined to generate updated structured analysis results. Specifically, the process involves: inputting the real-time gaze point sequence into the emotion state analysis model optimized by real-time cognitive feedback; the model performing feature extraction on the real-time gaze point sequence; and using a hierarchical extraction method with a preset hierarchical extraction threshold of feature intensity 0.2, where features below this threshold are discarded. On one hand, the model extracts visual gaze point sequence pattern features contained in the real-time gaze point sequence, specifically including the movement trajectory of the gaze point, the switching frequency between different attention distribution units, and the... Features such as the dwell time distribution of the unit and the movement speed of the gaze point are extracted. The switching frequency is calculated as the number of times the attention distribution unit is switched per unit time, and the movement speed is calculated as the distance between two adjacent gaze points divided by the time difference between the two gaze points. On the other hand, the physiological signal rhythm features associated with the real-time gaze point sequence are extracted. The physiological signals come from the physiological acquisition devices on the smart reading terminal, specifically the heart rate sensor and the skin conductance sensor. The collected physiological signals include the child's heart rate, skin conductance signals, etc. Features such as rhythm changes, fluctuation amplitude, and peak frequency of these physiological signals are extracted. The fluctuation amplitude is calculated as the maximum value of the physiological signal minus the minimum value of the physiological signal, and the peak frequency is calculated as the number of times the peak value of the physiological signal occurs per unit time.

[0085] Based on the extracted visual gaze point sequence pattern features and physiological signal rhythm features, a comprehensive analysis and calculation of the emotional state analysis model optimized by real-time cognitive feedback is performed. The calculation method for the cognitive attention quantification is as follows: the feature value corresponding to the visual gaze point sequence pattern feature is multiplied by the corresponding weight in the updated attention weight matrix, and then the feature value corresponding to the physiological signal rhythm feature is multiplied by the corresponding weight in the updated attention weight matrix. After obtaining the weighted sum, the weighted sum is compared with the updated emotional feature threshold. If the weighted sum is greater than or equal to the updated emotional feature threshold, the original weighted sum is retained; if the weighted sum is less than the updated emotional feature threshold, it is set to 0. The final cognitive attention metric is obtained through this calculation. The probability of emotional category affiliation is calculated by the model fusing the extracted two types of features and then calculating it through a preset softmax classification function. The specific implementation process is as follows: first, the feature output value of each emotional category is calculated, then the feature output value of each emotional category is used as the exponent of the exponential function to calculate the exponent value, and then the exponent value of each emotional category is divided by the sum of the exponent values ​​of all emotional categories to obtain the probability of affiliation to that emotional category. The emotional categories that children may be in include pleasure, boredom, curiosity, and irritability. The probability value of each emotional category is calculated in this way.

[0086] The emotional intensity level index is calculated by comprehensively quantifying the maximum probability of emotional category assignment, combined with the fluctuations in visual gaze point sequence pattern features and physiological signal rhythm features. Specifically, the maximum probability of emotional category assignment is multiplied by a feature fluctuation coefficient, which is the average of the physiological signal fluctuation amplitude and the visual gaze point movement speed. This product is then divided by a preset baseline coefficient of 0.5 to obtain the emotional intensity level index. Finally, the calculated cognitive attention quantification value, emotional category assignment probability, and emotional intensity level index are integrated and arranged in a fixed field order. First, the cognitive attention quantification value is presented, retaining three decimal places. Then, each emotional category and its corresponding assignment probability are presented in descending order of probability value, with each emotional category corresponding to one line, the emotional category name first, the assignment probability second, and retaining three decimal places. Finally, the emotional intensity level index is presented, retaining three decimal places. No additional formatting is required; the categories are simply arranged in this fixed order and combined to generate the updated structured analysis results.

[0087] The updated structured analysis results are input into the matching and mapping process to generate an updated AI picture book strategy parameter set. This updated set is then input into the picture book visual content generation process to adjust the parameters of the visual content to be generated, enabling real-time emotional interaction guidance. Specifically, this includes: determining the child's current dominant emotional category based on the emotional category attribution probability in the updated structured analysis results. The dominant emotional category is selected by choosing the category with the highest probability. If two or more emotional categories have the same probability value, the category with the highest preset priority is chosen. The preset priority order is curiosity, pleasure, boredom, and irritability, based on the child's reading experience and emotional guidance effect, prioritizing the cultivation of the child's reading interest. After determining the dominant emotional category, a one-to-one precise match is performed from a preset emotional guidance strategy library. The matching algorithm within the strategy library compares the dominant emotional category with the guidance strategies corresponding to various emotions in the library, identifying the most suitable match for the current dominant emotional category. This database of basic emotional guidance strategies is built upon the emotional cognitive patterns and picture book reading guidance needs of children aged 3 to 8, through extensive experiments on children's emotional cognition. Each emotional category in the database corresponds to a set of exclusive basic emotional guidance strategies. Each set of basic emotional guidance strategies explicitly includes core content such as visual presentation parameters, content rhythm parameters, and interaction method parameters. Based on the cognitive attention measurement values ​​in the updated structured analysis results, the visual presentation parameters in the basic emotional guidance strategies are adjusted. These visual presentation parameters include baseline values ​​for visual element complexity, color saturation, dynamic effects, and text-image ratio. Specifically, the adjustment method is as follows: the baseline value for visual element complexity is added to the cognitive attention measurement value multiplied by 0.2 to obtain the adjusted visual element complexity parameter; the baseline value for color saturation is added to the cognitive attention measurement value multiplied by 0.1 to obtain the adjusted color saturation parameter; the baseline value for dynamic effects is added to the cognitive attention measurement value multiplied by 0.1 to obtain the adjusted dynamic effects parameter; and the baseline value for text-image ratio is added to the cognitive attention measurement value multiplied by 0.15 to obtain the adjusted text-image ratio parameter. After all visual presentation parameters are adjusted, the adjusted visual presentation parameters are integrated with other core parameters in the basic emotional guidance strategy, including content rhythm parameters and interaction method parameters. During integration, the parameters are classified according to their functional attributes to ensure that the parameters are arranged in an orderly and clear manner, thereby generating a primary strategy parameter set.

[0088] Based on the emotional intensity level index in the updated structured analysis results, the primary strategy parameter set is intensity modulated. The modulation rules are pre-defined to achieve precise matching between parameters and emotional intensity. Specifically, each parameter in the primary strategy parameter set is multiplied by the emotional intensity level index and then added to its corresponding baseline offset. The baseline offset for visual element complexity is 0.02, for color saturation it is 0.01, for dynamic effects it is 0.01, for text-image ratio it is 0.02, and for content rhythm and interaction methods it is 0. This ensures that the modulation effect of different types of parameters aligns with the emotional guidance requirements. During the modulation process, each parameter in the primary strategy parameter set is calculated and modulated one by one. After the calculation, the range of each modulated parameter is checked to ensure that each modulated parameter is within the preset reasonable range, avoiding deviations in subsequent visual content generation caused by parameter anomalies. After all parameters are modulated, all modulated parameters are integrated, duplicate parameters are removed, and they are arranged according to the preset parameter classification order to obtain the complete strategy parameter set.

[0089] The complete set of strategy parameters is categorized and organized according to their functional attributes, specifically into four categories: visual element control parameters, narrative rhythm control parameters, emotional guidance control parameters, and interactive effect control parameters. Real-time verification is performed during the categorization process, checking the functional attributes of each parameter one by one to ensure accurate classification and prevent errors or omissions. After categorization, the parameters are arranged in a preset order: visual element control parameters, narrative rhythm control parameters, emotional guidance control parameters, and interactive effect control parameters. This order is based on the priority of parameters in the generation of picture book visual content, facilitating subsequent parameter parsing and use. Each parameter undergoes secondary verification during the arrangement process to check if its value is within a preset reasonable range and to ensure consistent terminology in parameter names, avoiding confusion and ensuring parameter standardization. Finally, all the arranged parameters are integrated according to a preset format to generate an updated AI picture book strategy parameter set.

[0090] The updated AI picture book strategy parameter set is input into the picture book visual content generation process. First, the visual element complexity adjustment coefficient, narrative rhythm control parameter, emotional guidance strategy identifier, dynamic effect intensity parameter, and color saturation adjustment parameter are parsed from the AI ​​picture book strategy parameter set. During parsing, a preset key-value pair matching parsing algorithm is used. The algorithm is first initialized, and the target key names of the five core parameters are preset. At the same time, the key-value comparison threshold is set to 0.95. After initialization, the algorithm reads all key-value pair data in the AI ​​picture book strategy parameter set. After splitting according to the preset byte splitting rules, the key name strings are processed to remove spaces and unify uppercase and lowercase. Then, according to the preset parsing order, the similarity between the processed key name strings and the target key names of the five core parameters is calculated. Based on the similarity results, the corresponding numerical or identifier content is extracted as the initial extraction value of the core parameters. Then, the initial extraction value is converted into a standard format according to the preset data format conversion rules. Finally, range or legality verification is performed to ensure that the parameters are valid and usable.

[0091] Based on the visual element complexity adjustment coefficient and narrative rhythm control parameters, a preset picture book material library is invoked to select matching character images, background scenes, and decorative objects. This preset picture book material library is constructed based on the cognitive characteristics and reading preferences of children aged 3 to 8. The materials in the library are divided into three levels according to visual element complexity: simple, moderate, and rich; and into three categories according to narrative rhythm: soothing, moderate, and fast. After invoking, the complexity level and appropriate rhythm of the materials are determined according to the parameters. Matching materials are selected and combined according to the picture book narrative logic to generate the basic visual composition of the picture book page. At the same time, a preset page layout rule library is invoked. This rule library is constructed based on the visual cognitive patterns of children aged 3 to 8. Based on the visual element complexity adjustment coefficient and narrative rhythm control parameters, matching text and image arrangement and area ratio are selected to generate the framework layout of the picture book page.

[0092] Based on the emotional guidance strategy identifiers and basic visual components, position offset reference values, angle deflection reference values, and shape scaling reference values ​​corresponding to each component are extracted from a pre-set emotional action mapping rule base. These serve as initial position adjustment rules, initial angle adjustment rules, and initial shape adjustment rules. This rule base was constructed through numerous emotional cognition experiments with children aged 3 to 8. Different emotional guidance strategy identifiers in the base correspond to different visual element adjustment reference values. Based on the dynamic effect intensity parameter, proportional calculations are performed on the position offset reference values, angle deflection reference values, and shape scaling reference values. The position movement is calculated by multiplying the position offset reference value of each visual element by the dynamic effect intensity parameter, and the angle rotation is calculated by multiplying the angle deflection reference value of each visual element by the dynamic effect intensity parameter. The dynamic effect intensity parameter and the form adaptation amount are calculated by multiplying the form scaling baseline value of each visual element by the dynamic effect intensity parameter and adding the preset baseline scaling value of 1.0. After the calculation is completed, a range check is performed to ensure that the adjustment amount is reasonable. According to the color saturation adjustment parameter, the original color values ​​of each element in the basic visual composition are superimposed. The original color value of each visual element is added to the color saturation adjustment parameter and then subtracted from the preset baseline correction value of 0.5. Two decimal places are retained during the calculation. After the calculation is completed, a range check is performed to control the adjusted color saturation value within a reasonable range of 0.3 to 0.9. After the color saturation of each visual element is adjusted, a comparison check is performed to ensure that the adjusted color is coordinated and unified with the colors of other elements.

[0093] Position displacement, angular rotation, and shape adaptation are applied to the corresponding character image, background scene, and decorative objects in the basic visual composition, respectively. Deformation is performed sequentially in the order of translation, rotation, and scaling. During translation, the original coordinates of the element are added with the corresponding position displacement. During rotation, the element's center coordinates are used as the rotation center point, and the element's center coordinates are used as the scaling center point, scaling with the corresponding shape adaptation. After these three transformation processes are completed, each visual element is validated to ensure there are no abnormal deviations, thus obtaining the deformable adaptation. The visual elements are then placed in the picture book page frame according to the order of background scene, character image, decorative object, and text elements. The adjusted color saturation values ​​are assigned to each element after placement. Each element is verified to ensure normal presentation. After all elements are placed, an overall visual verification is performed to check whether the page layout is reasonable, whether the element proportions are coordinated, whether the colors are unified, whether there are any abnormal deviations, and whether the dynamic effects are adapted. After the verification is passed, new picture book visual content is generated. Through this process, the parameters of the picture book visual content to be generated are continuously adjusted to achieve real-time emotional interaction guidance.

[0094] In this embodiment, an updated set of AI picture book strategy parameters is generated based on the updated structured analysis results. This ensures that the guidance strategy is highly adapted to the child's real-time state. By inputting this set of parameters into the picture book visual content generation process, the picture book visual content parameters can be adjusted in real time, realizing personalized real-time emotional interaction guidance, and allowing the picture book content to dynamically match the child's reading state.

[0095] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the system as described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.

[0096] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An AI-powered emotional interaction guidance system for children's picture book reading, characterized in that, include: The data acquisition module is used to collect data streams of children's reading behavior. The computation module is used to construct the attention distribution plane based on the reading behavior data stream; Based on the attention distribution plane, identify the areas where children's attention is focused, and calculate the minimum bounding rectangle of the areas to generate multiple attention distribution units; based on the distribution density and duration of reading behavior data streams within each attention distribution unit, calculate the cognitive distribution feature value. The optimization module is used to optimize the emotional state analysis model based on cognitive distribution feature values. The optimized model is used to process the reading behavior data stream to obtain structured analysis results that characterize children's cognitive and emotional states. The structured analysis results include cognitive attention quantification values, emotional category attribution probability, and emotional intensity level index. The matching module is used to perform matching and mapping based on the structured analysis results to obtain a set of AI picture book strategy parameters, including: matching a preset emotional guidance strategy library according to the probability of emotional category affiliation to obtain the basic emotional guidance strategy most suitable for the current emotional category; adjusting the visual presentation parameters in the basic emotional guidance strategy according to the cognitive attention metric value to generate a primary strategy parameter set; modulating the intensity of the primary strategy parameter set according to the emotional intensity level index to obtain a complete strategy parameter set; and integrating the complete strategy parameter set to generate the AI ​​picture book strategy parameter set. The processing module is used to generate the basic visual composition and framework layout of the picture book page based on the AI ​​picture book strategy parameter set; determine the position, angle and shape adjustment rules of each component of the basic visual composition based on the emotional guidance parameters in the AI ​​picture book strategy parameter set; and perform position movement, angle rotation and shape adaptation on the basic visual composition and framework layout according to the shape adjustment rules to generate picture book visual content. The update module is used to collect the dynamic visual attention data stream of readers in real time during the presentation of picture book visual content, and convert it into updated cognitive distribution feature values, which are then fed back to the emotional state analysis model to dynamically adjust the generation of subsequent picture book visual content and achieve emotional interaction guidance.

2. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 1, characterized in that, Based on reading behavior data stream, an attention distribution plane is constructed; according to the attention distribution plane, the children's attention concentration areas are identified, and the minimum bounding rectangle of the concentration areas is calculated to generate multiple attention distribution units; Based on the distribution density and duration of the reading behavior data stream within each attention distribution unit, cognitive distribution feature values ​​are calculated, including: Analyze the reading behavior data stream, extract the visual fixation point coordinate sequence, and construct an attention distribution plane based on the position of all visual fixations on the picture book page. On the attention distribution plane, visual fixation points are analyzed to identify multiple fixation point groups that form clusters in space, and the area covered by the fixation point groups is determined as the attention concentration area. For each attention-focused region, calculate the minimum bounding rectangle that covers all visual fixations within the corresponding attention-focused region, and generate multiple attention distribution units; Based on the attention distribution unit, the number of visual fixations falling within the attention distribution unit is counted, and the unit distribution density is calculated by combining the area of ​​the attention distribution unit; the average duration of visual fixations falling within the attention distribution unit is calculated to obtain the unit persistence intensity. The cognitive contribution of a unit is calculated by fusing the unit distribution density and the unit persistence intensity; the cognitive distribution characteristic value is calculated by integrating the cognitive contribution of all units.

3. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 2, characterized in that, For each attention-focused region, the minimum bounding rectangle covering all visual fixations within that region is calculated, generating multiple attention distribution units, including: Based on the attention focus area, extract the gaze coordinates of all visual gaze points that constitute the attention focus area; Based on all fixation point coordinates, find the minimum and maximum values ​​of the x-coordinate and the y-coordinate respectively. Determine the coordinates of the top-left vertex of the minimum bounding rectangle based on the minimum x-coordinate and the maximum y-coordinate; determine the coordinates of the bottom-right vertex of the minimum bounding rectangle based on the maximum x-coordinate and the minimum y-coordinate. Based on the coordinates of the top left and bottom right vertices, a minimum bounding rectangle is generated that completely encloses the gaze point coordinates and whose edges are parallel to the coordinate axes of the picture book page; based on the minimum bounding rectangle, the attention distribution unit is obtained.

4. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 3, characterized in that, Based on the attention distribution unit, the number of visual fixations falling within the attention distribution unit is counted, and the unit distribution density is calculated by combining the area of ​​the attention distribution unit. Calculate the mean duration of visual fixation points falling within the attention distribution unit to obtain the unit persistence intensity, including: Based on the attention distribution unit, all visual fixation points located within the boundary range of each attention distribution unit are obtained, generating a subset of visual fixation points for each attention distribution unit; and the duration corresponding to each visual fixation point in the subset of visual fixation points is obtained, resulting in a subset of durations. Based on the subset of visual fixation points, count the total number of visual fixation points in the subset; divide the total number of visual fixation points by the area of ​​the corresponding attention distribution unit to obtain the unit distribution density; Based on the duration subset, sum all duration values ​​in the duration subset; divide the sum by the number of data in the duration subset to obtain the cell duration intensity.

5. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 4, characterized in that, The emotional state analysis model was optimized based on cognitive distribution eigenvalues. The optimized model was then used to process reading behavior data streams, yielding structured analysis results characterizing children's cognitive and emotional states, including: The cognitive distribution feature values ​​are input into a preset model parameter mapping function; the model parameter mapping function outputs the emotional feature threshold adjustment and attention weight adjustment of the emotional state analysis model. Based on the adjustment of the emotional feature threshold and the adjustment of the attention weight, the original emotional feature threshold and the original attention weight matrix stored in the emotional state analysis model are updated to obtain the optimized emotional state analysis model. The reading behavior data stream is input into the optimized emotion state analysis model to extract the visual fixation point sequence pattern features and physiological signal rhythm features contained in the reading behavior data stream; Based on the visual gaze point sequence pattern characteristics and physiological signal rhythm characteristics, the optimized emotion state analysis model calculates the cognitive attention quantification value, the probability of emotion category attribution, and the emotion intensity level index. The combined cognitive focus quantification, probability of emotional category attribution, and emotional intensity level index generate structured analysis results.

6. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 5, characterized in that, Based on the adjustment of the sentiment feature threshold and the attention weight, the original sentiment feature threshold and the original attention weight matrix stored internally in the sentiment state analysis model are updated to obtain the optimized sentiment state analysis model, including: Based on the emotional feature threshold adjustment amount, the preset original emotional feature thresholds in the emotional state analysis model are read; the original emotional feature thresholds are added to the emotional feature threshold adjustment amount to generate the updated emotional feature thresholds; based on the attention weight adjustment amount, the preset original attention weight matrix in the emotional state analysis model is read. Each element in the original attention weight matrix is ​​added to the corresponding adjustment value in the attention weight adjustment matrix to generate the updated attention weight matrix. The updated sentiment feature thresholds and the updated attention weight matrix are written into the corresponding storage location of the sentiment state analysis model to complete the model parameter update and obtain the optimized sentiment state analysis model.

7. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 6, characterized in that, The basic visual composition and framework layout of the picture book page are generated based on the AI ​​picture book strategy parameter set; based on the emotional guidance parameters in the AI ​​picture book strategy parameter set, the position, angle, and shape adjustment rules of each component of the basic visual composition are determined; according to the shape adjustment rules, the basic visual composition and framework layout are moved, rotated, and adapted to generate the picture book visual content, including: The following parameters were obtained from the AI ​​picture book strategy parameter set: visual element complexity adjustment coefficient, narrative rhythm control parameter, emotional guidance strategy identifier, dynamic effect intensity parameter, and color saturation adjustment parameter. Based on the complexity adjustment coefficient of visual elements and the narrative rhythm control parameters, the preset picture book material library is called to select matching character images, background scenes and decorative objects, and the basic visual composition of the picture book page is generated. The preset page layout rule library is called to select matching text and image arrangement and area ratio, and the framework layout of the picture book page is generated. Based on the emotional guidance strategy identifier and basic visual composition, the position offset reference value, angle deflection reference value and shape scaling reference value corresponding to each component are extracted from the preset emotional action mapping rule library and used as the initial position adjustment rule, initial angle adjustment rule and initial shape adjustment rule. Based on the dynamic effect intensity parameters, proportional calculations are performed on the position offset reference value, angle deflection reference value and shape scaling reference value to generate the final position movement amount, angle rotation amount and shape adaptation amount of each component. Based on the color saturation adjustment parameters, the original color values ​​of each element in the basic visual composition are superimposed to generate the adjusted color saturation values. The position shift, angle rotation, and shape adaptation are applied to the corresponding character image, background scene, and decorative object in the basic visual composition, respectively, to complete the translation, rotation, and scaling transformations in the coordinate system of the picture book page, and obtain the visual composition elements after shape adaptation. The deformed and adapted visual elements are placed into the picture book page frame in sequence, and the adjusted color saturation values ​​are assigned to generate the picture book visual content.

8. The AI-powered emotional interaction guidance system for children's picture book reading according to claim 7, characterized in that, During the presentation of picture book visual content, the dynamic visual attention data stream of readers is collected in real time and converted into updated cognitive distribution feature values, which are then fed back to the emotional state analysis model to dynamically adjust the generation of subsequent picture book visual content and achieve emotional interaction guidance, including: The system collects dynamic visual attention data streams from readers during the presentation of picture book visual content in real time, and generates a real-time gaze point sequence. Based on the position coordinates of all real-time gaze points in the real-time gaze point sequence on the current picture book page, a real-time attention distribution plane is constructed. Analyze all real-time gaze points on the real-time attention distribution plane to obtain multiple real-time gaze point groups. Determine the area covered by each real-time gaze point group as the real-time attention gathering region and generate a set of real-time attention gathering regions. For each real-time attention gathering region, calculate the minimum bounding rectangle of all real-time gaze points within the complete coverage area, and determine each minimum bounding rectangle as a real-time attention distribution unit to generate a set of real-time attention distribution units. Based on the set of real-time attention distribution units and the sequence of real-time gaze points, the number of real-time gaze points and the duration of gaze in each real-time attention distribution unit are counted, and the real-time unit distribution density and the real-time unit duration intensity are calculated to generate the set of real-time unit distribution density and the set of real-time unit duration intensity. The real-time unit distribution density set and the corresponding values ​​belonging to the same real-time attention distribution unit in the real-time unit continuous intensity set are weighted and fused to calculate the real-time unit cognitive contribution, and a real-time unit cognitive contribution set is generated. The real-time cognitive contribution of each unit is summed and the summation result is normalized to obtain the updated cognitive distribution feature value. The updated cognitive distribution feature value is then input into the emotional state analysis model, and the parameters of the emotional feature threshold and attention weight matrix stored in the model are updated to obtain the emotional state analysis model optimized by real-time cognitive feedback. The real-time gaze sequence is input into the emotion state analysis model optimized by real-time cognitive feedback. The updated cognitive attention quantification, the updated emotion category classification probability, and the updated emotion intensity level index are obtained through processing. The updated structured analysis results are then generated by combining these results. The updated structured analysis results are input into the matching and mapping process to generate an updated set of AI picture book strategy parameters. The updated set of AI picture book strategy parameters is then input into the picture book visual content generation process to adjust the parameters of the picture book visual content to be generated, thereby achieving real-time emotional interaction guidance.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the system as described in any one of claims 1 to 8.