A puzzle interaction control method, device and equipment of an intelligent point-and-read pen and a medium
Patent Information
- Application Number
- CN202611003468.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-07
AI Technical Summary
[0004]本发明的主要目的在于提供一种智能点读笔的拼图交互控制方法、装置、设备及介质,旨在解决现有技术中点读笔拼图交互模式单一、自适应性和多样性不够的技术问题
[0009] Beneficial Effects: This invention discloses a puzzle interaction control method, device, equipment, and medium for an intelligent reading pen. Compared to existing technologies, this invention, when a puzzle piece recognition event is triggered, acquires the identification information of the target puzzle piece; acquires the theme resource data corresponding to the target puzzle piece based on the identification information; acquires the user profile data and puzzle context data of the current user; determines the current interaction mode based on the received user input data and the puzzle context data; matches a corresponding interaction template according to the current interaction mode; injects the user profile data and puzzle context data as constraint parameters into the interaction template; and generates a corresponding initial response based on the theme resource data; adjusts the initial response based on the user's cognitive level in the user profile data; and generates and outputs the final response. By integrating puzzle piece recognition with user profiles and context for interaction mode switching, and dynamically adjusting the response content based on cognitive level, the same puzzle piece can output differentiated content according to different situations, realizing dynamic interaction of puzzle reading and improving the interactive adaptability and diversity of puzzle teaching aids.
Smart Images

Figure CN122507290B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, and medium for interactive puzzle control of an intelligent reading pen. Background Technology
[0002] Smart teaching tools such as reading pens can be used in conjunction with puzzles and other media. By lightly touching the puzzle with the reading pen, corresponding interactive content is output, enabling children to learn while playing. Therefore, they are widely used in early childhood education.
[0003] Existing reading pens or educational toys typically only play pre-recorded audio based on code points or labels after the puzzle is touched. The output content is fixed and has a single mode, making it difficult to form a dynamic interaction with the user's actual puzzle situation and the characteristics of different users, thus reducing the adaptability and diversity of reading pen puzzle interaction. Summary of the Invention
[0004] The main objective of this invention is to provide a method, device, equipment, and medium for controlling the puzzle interaction of an intelligent reading pen, aiming to solve the technical problems of the existing reading pen's single puzzle interaction mode, insufficient adaptability, and lack of diversity.
[0005] The technical solution of the present invention is as follows: The first aspect of this invention provides a puzzle interaction control method for an intelligent reading pen, comprising: When the puzzle piece recognition event is triggered, obtain the identification information of the target puzzle piece; Obtain the theme resource data corresponding to the target puzzle piece based on the identification information; Obtain the current user's user profile data and puzzle context data, and determine the current interaction mode based on the received user input data and the puzzle context data; Match the corresponding interaction template according to the current interaction mode, inject the user profile data and puzzle context data as constraint parameters into the interaction template, and generate the corresponding initial response in combination with the topic resource data; The initial response is adjusted based on the user's cognitive level in the user profile data to generate and output the final response.
[0006] A second aspect of the present invention provides a puzzle interaction control device for a smart reading pen, comprising: The puzzle recognition module is used to obtain the identification information of the target puzzle piece when the puzzle piece recognition event is triggered; The resource acquisition module is used to acquire the theme resource data corresponding to the target puzzle piece based on the identification information; The interaction mode switching module is used to obtain the current user's user profile data and puzzle context data, and determine the current interaction mode based on the received user input data and the puzzle context data; The response generation module is used to match the corresponding interaction template according to the current interaction mode, inject the user profile data and puzzle context data as constraint parameters into the interaction template, and generate the corresponding initial response in combination with the topic resource data; The cognitive adjustment module is used to adjust the initial response based on the user's cognitive level in the user profile data, and generate and output the final response.
[0007] A third aspect of the present invention provides a computer device including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the aforementioned puzzle interaction control method of the smart reading pen.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the above-described puzzle interaction control method for a smart reading pen.
[0009] Beneficial Effects: This invention discloses a puzzle interaction control method, device, equipment, and medium for an intelligent reading pen. Compared to existing technologies, this invention, when a puzzle piece recognition event is triggered, acquires the identification information of the target puzzle piece; acquires the theme resource data corresponding to the target puzzle piece based on the identification information; acquires the user profile data and puzzle context data of the current user; determines the current interaction mode based on the received user input data and the puzzle context data; matches a corresponding interaction template according to the current interaction mode; injects the user profile data and puzzle context data as constraint parameters into the interaction template; and generates a corresponding initial response based on the theme resource data; adjusts the initial response based on the user's cognitive level in the user profile data; and generates and outputs the final response. By integrating puzzle piece recognition with user profiles and context for interaction mode switching, and dynamically adjusting the response content based on cognitive level, the same puzzle piece can output differentiated content according to different situations, realizing dynamic interaction of puzzle reading and improving the interactive adaptability and diversity of puzzle teaching aids. Attached Figure Description
[0010] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 A schematic diagram of an application environment for the puzzle interaction control method of the intelligent reading pen provided in an embodiment of the present invention; Figure 2 A flowchart of a puzzle interaction control method for an intelligent reading pen provided in an embodiment of the present invention; Figure 3 A schematic diagram of the functional modules of the puzzle interaction control device for the intelligent reading pen provided in an embodiment of the present invention; Figure 4 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.
[0013] The puzzle interaction control method for the intelligent reading pen provided in this embodiment of the invention can be applied to, for example... Figure 1 In the interactive scenario shown, where the smart reading pen is paired with the puzzle device, there are a terminal device 101, a smart reading pen 102, a network 103, and a server 104. The network 103 serves as a medium to provide a communication link between the terminal device 101, the smart reading pen 102, and the server 104. The network 103 can include various connection types, such as wired and / or wireless communication links (e.g., Bluetooth, Wi-Fi, NFC, etc.).
[0014] Users can use terminal device 101 and smart reading pen 102 to interact with server 104 via network 103 to receive or send messages, etc. A client supporting this interaction method can be installed on terminal device 101, allowing users to log in to view and edit the interaction settings of smart reading pen 102, view puzzle interaction records, etc. Terminal device 101 can be various electronic devices with a display screen and web browsing support, including but not limited to smartphones, tablets, and desktop computers.
[0015] The smart reading pen 102 is used in conjunction with the puzzle piece 105. When the user places the smart reading pen 102 on the puzzle piece 105, it can automatically recognize the identifiable labels on the surface of the puzzle piece, thereby triggering a recognition event and outputting corresponding interactive content in different interactive states. Simultaneously, the user can press the button on the smart reading pen 102 to input voice through its built-in microphone for intelligent voice dialogue, achieving accurate transmission of interactive needs. The smart reading pen 102 integrates motion sensing, near-field communication recognition, voice acquisition, and feedback output functions, enabling it to sense the user's touch, press, and other operations in real time and associate them as corresponding interactive events.
[0016] Server 104 can be a server providing various services, such as analyzing and processing data like puzzle information recognized by the smart reading pen 102 and user-inputted voice commands, generating AI interactive content, and feeding it back to the smart reading pen 102's backend server (this is just an example). Server 104 can parse and process the received puzzle recognition data and user interaction commands, generate adapted interactive feedback content, and then transmit it to the smart reading pen 102 via network 103, where the smart reading pen 102 provides feedback to the user in the form of voice, light, etc. Server 104 can be a cloud server, a distributed system server, or a server integrated with blockchain.
[0017] It should be understood that the number of terminal devices 101, smart reading pens 102, networks 103, and servers 104 mentioned above is merely illustrative. Depending on the implementation needs, there can be any number of terminal devices 101, smart reading pens 102, networks 103, and servers 104.
[0018] like Figure 2 As shown, the puzzle interaction control method of the intelligent reading pen provided in this embodiment of the invention specifically includes the following steps: S201. When the puzzle piece recognition event is triggered, obtain the identification information of the target puzzle piece.
[0019] In this embodiment, the puzzle piece recognition event is triggered when the user presses, touches, or approaches the puzzle piece with an identification tag (such as an RFID tag or visual identification code) on its surface. After the puzzle piece recognition event is triggered, the identification information of the target puzzle piece is obtained through a near-field communication mechanism. Specifically, the pen terminal has a built-in near-field communication module (such as an NFC or RFID reader), and each puzzle piece contains an electronic tag that stores identification information, resource indexes, version information, topic numbers, or other data that can trigger content retrieval. When the pen tip presses, touches, or approaches the puzzle piece, the identification information of the target puzzle piece in the electronic tag can be read.
[0020] Preferably, during the reading of the identification information, the pressing duration, number of presses, and pen tip status can be further detected. This action data can be used as part of the subsequent user input data to comprehensively determine the user's point-reading interaction intent. By using the identification information as the basis for dynamic interaction rather than just static tags used to index fixed audio, a foundation is provided for subsequent dynamic puzzle point-reading interaction that combines user input and user profile information.
[0021] For example, in a children's puzzle scenario, a child presses a reading pen on a puzzle piece marked with a "lion" pattern. The reading pen reads the puzzle piece's identification information as "SAVANNA_THEME_003" via NFC. This identification information also carries the theme number "SAVANNA" and the piece number "003", thereby confirming that the current interactive object is a lion character puzzle piece under the African savanna theme, and preparing to enter the subsequent resource location and interaction determination process.
[0022] S202. Obtain the theme resource data corresponding to the target puzzle piece based on the identification information.
[0023] In this embodiment, based on the identification information of the target puzzle piece, the corresponding theme resource data is retrieved. This theme resource data is a structured content collection bound to the identification information. Unlike traditional point-and-read recognition, this theme resource data is not a pre-stored single audio file, but contains multi-dimensional data such as story script fragments, knowledge point entries, character attribute descriptions, scene associations, image material indexes, semantic tags, and cloud resource addresses. This provides rich materials for the subsequent generation of response content, thereby improving the diversity of the generated content.
[0024] Specifically, the topic resource data can be obtained through a collaborative approach between local and cloud platforms. First, a query is performed in the local content library, using the topic number and unique block number from the identification information as the joint query key to retrieve the corresponding topic resource data. If the local cache is hit and the version verification passes, the local topic resource package is directly loaded. If the local cache misses, the version expires, or the resource package is missing data, the cloud server is requested to download supplementary data using the cloud resource address or resource index carried in the identification information. This achieves collaborative acquisition of local basic resources and cloud-enhanced resources. In this embodiment, by acquiring multi-dimensional topic resource data, it is possible to organize narrative or Q&A content around the semantic role (such as animal, scene, prop, etc.) of the current puzzle piece, rather than simply playing pre-recorded fixed audio. This supports outputting differentiated content for the same puzzle piece in different interaction modes, improving interaction flexibility.
[0025] For example, based on the identifier "SAVANNA_THEME_003", a "African Savannah" themed resource pack is loaded from the local database. This resource pack includes the background story of the lion character, dietary knowledge, relationship data with other savannah animals such as zebras and giraffes, and the expected position coordinates of the puzzle piece in the whole puzzle. If the search finds that the local resource pack is missing some open-ended question and answer materials, it will automatically request them from the cloud through the resource index.
[0026] S203. Obtain the current user profile data and puzzle context data, and determine the current interaction mode based on the received user input data and the puzzle context data.
[0027] In this embodiment, user profile data and puzzle context data are read from a local cache database. The user profile data is structured data describing the user's cognitive abilities and interaction preferences, including at least parameters such as user cognitive level, age group, grade level, language ability tags, parent-set goals, historical frequency of requests for help, and number of consecutive incorrect answers. The puzzle context data is structured data describing the current puzzle progress status, including at least the current puzzle theme, distribution of completed areas, semantic tags of target puzzle pieces, relationships between adjacent pieces, overall puzzle structure, current progress percentage, and identifier of the most recently interacted piece.
[0028] Furthermore, user input data is received by collecting pen tip movements and real-time monitoring of ambient speech. This user input data includes speech data collected via the microphone and action data collected via the pen tip sensors, such as press duration, number of presses, pause duration, and re-trigger interval. After parsing, this user input data forms corresponding action signals and speech semantics. Based on these action signals and speech semantics, combined with user profiles and puzzle context, interaction pattern recognition is performed to determine the current interaction mode. Specific interaction modes can include location assistance mode, error correction confirmation mode, knowledge question-and-answer mode, storytelling mode, repeated recognition switching mode, cross-block related dialogue mode, and default storytelling or free chat mode, etc. By making interaction mode decisions based on user personalization characteristics, puzzle context, and real-time user operations, the same puzzle piece can trigger differentiated interaction paths under different users, different progress levels, and different operations, thus solving the problems of existing reading pen products having single output and lacking personalized interaction perception.
[0029] For example, when the user profile data read includes cognitive level L2 (standard interaction level), age 5, language ability label "intermediate", parent-set goal "science popularization and hands-on ability", and 0 consecutive incorrect answers; the puzzle context data shows the current progress is 60%, the most recently interacted piece is the "zebra piece", the current theme is still "African savanna", and the candidate empty areas are the two positions in the lower right corner. The user input data includes a single short press action without a voice question. At this time, the comprehensive analysis shows that the current situation meets the criteria for initial recognition without accompanying question, thus determining that the current interaction mode is storytelling mode, and entering storytelling mode to conduct subsequent interactions with the user.
[0030] S204. Match the corresponding interaction template according to the current interaction mode, inject the user profile data and puzzle context data as constraint parameters into the interaction template, and generate the corresponding initial response in combination with the topic resource data.
[0031] In this embodiment, the interaction template is a predefined structured response generation framework. Different interaction modes correspond to different template sets. The basic templates include at least storytelling templates, knowledge Q&A templates, location prompt templates, error correction feedback templates, and encouragement feedback templates, etc. Based on the confirmed current interaction mode, a matching interaction template is retrieved from the preset template library. Then, user profile data and puzzle context data are injected as constraint parameters into the currently matched interaction template, transforming the interaction template from a general framework into a personalized generation task tailored to the current user and the current scenario.
[0032] The text generation task is executed based on an interactive template with injected constraint parameters. Combined with acquired topic resource data, corresponding initial responses are generated. Specifically, the output generation can be performed by a local rule engine, a local knowledge base, or a cloud-based generation model. The actual execution entity can be flexibly switched based on current network conditions, task replication levels, and local resource availability. For example, local priority processing scenarios include initial storytelling, standard knowledge point announcements, cached fixed questions and answers, low-complexity location hints, and scenarios with unavailable networks or low latency requirements. Cloud-based enhanced processing scenarios include open-ended questions, cross-block related questions and answers, multi-turn contextual reasoning, questions with non-standard user expressions requiring semantic correction, and answers requiring dynamic rewriting based on user profiles. Ideally, local hits on cached and local rule answers can be attempted first. If a hit fails, the hit confidence is insufficient, or the question complexity is high, cloud generation is then invoked. This ensures basic interaction can be completed even without a network, and enhanced question answering and complex reasoning can be completed when the network is good, balancing interaction response speed and generation capabilities.
[0033] For example, in the current interaction mode, a "storytelling template" is matched. The corresponding generation task is obtained by using "lion" as the character name constraint, "African savanna" as the background constraint, cognitive level L2 as the sentence length constraint, and the most recent interaction block "zebra" as the associated character constraint. The relevant story fragment is cached locally, so the generation task is executed directly on the local machine in combination with the theme resource data to generate the corresponding initial response: "The lion is the king of the savanna, and it has a golden mane. You just met the zebra, and it lives on the same savanna as the lion." S205. Adjust the initial response based on the user's cognitive level in the user profile data to generate and output the final response.
[0034] In this embodiment, since the local rule engine or cloud model that generates the initial response is trained based on comprehensive data from different users, although it can generate accurate response content in most cases, due to the different cognition of different users, there are obvious individual differences in their vocabulary, knowledge reserves, and reasoning abilities. Although the content of the initial response is accurate, there may be a situation of cognitive mismatch. For example, the content of the initial response may be too simple and childish for older children, or difficult for younger children to understand.
[0035] Therefore, this embodiment does not directly output the initial response to the user after it is generated. Instead, it adjusts the initial response based on the user's cognitive level in the user profile data. Specifically, it dynamically adjusts and adapts the initial response's sentence length, vocabulary difficulty, proportion of abstract concepts, openness of questions and answers, granularity of location prompts, and follow-up questioning strategies, thereby generating and outputting the final response in the form of voice, text, etc. This ensures that the final output content meets the user's comprehension ability and learning needs. While dynamically switching the interaction mode, it also achieves personalized cognitive adaptation of the output content, avoiding the situation of traditional reading pens having a single interaction mode, fixed output content, and inability to adapt to different children's cognitive levels. This improves the interactive adaptability and diversity of puzzle reading teaching aids.
[0036] For example, if the current user's cognitive level is L2, the corresponding difficulty indicator Df is set to "Medium", the help indicator Hf to "Location prompts are not applicable in the current scenario" (because the current mode is storytelling), and the follow-up question indicator Qf to "Add an open-ended follow-up question". After cognitive matching adjustment, the initial response maintains the original sentence length and vocabulary difficulty (meeting the L2 requirements), and a follow-up question is added at the end: "What do you think lions like to do most on the savanna?", thus generating the final response: "The lion is the king of the savanna, with a golden mane. You just met the zebra, which lives on the same savanna as the lion. What do you think lions like to do most on the savanna?". This final response is output through the speaker of the reading pen, prompting the user to interact again.
[0037] Specifically, user cognitive level is obtained through the following steps: Obtain basic profile parameters and interaction behavior parameters; The basic profile parameters and interaction behavior parameters are normalized, and then weighted and summed according to preset weights to obtain a comprehensive cognitive score. The comprehensive cognitive score is mapped to the corresponding user cognitive level according to a preset level range table.
[0038] In this embodiment, the user's cognitive level is obtained by comprehensively evaluating the user's basic profile parameters and interaction behavior parameters. The basic profile parameters are static attributes pre-entered by the user, including at least age group A, grade G, language ability L, and parent-set goals T. The interaction behavior parameters are dynamic behavior data collected and statistically analyzed in historical conversations, including at least the accuracy rate R of the most recent N rounds of questions, average response time D, number of consecutive requests for help H, number of consecutive incorrect answers E, and self-completion rate P.
[0039] First, normalize each parameter and map it to a preset range such as [0,1] to unify the parameter values of different dimensions. For example, age group A is mapped to a 0-1 value according to the preset age range; grade G is mapped to a 0-1 value according to the educational stage; language ability L is mapped to a 0-1 value according to the language assessment level; parent-set goal T is mapped to a 0-1 value according to the goal difficulty; the correct answer rate R in the most recent N rounds of questions is directly used as a 0-1 value; the average response time D is mapped to a 0-1 value according to the preset time range (the shorter the time, the higher the value); the number of consecutive requests for help H is mapped to a 0-1 value according to the preset number range (the fewer the number of times, the higher the value); the number of consecutive incorrect answers E is mapped to a 0-1 value according to the preset number range (the fewer the number of times, the higher the value); and the self-completion rate P is directly used as a 0-1 value. Then, a linear weighted summation is performed according to the preset weight configuration to calculate the comprehensive cognitive score: C = w1A + w2G + w3L + w4T + w5R + w6P - w7D - w8H - w9E, where C is the comprehensive cognitive score, and w1 to w9 are the weight values. Age group (A), grade (G), language ability (L), parent-set goals (T), correct answer rate (R), and independent completion rate (P) are positive factors, and therefore the summation sign is positive, indicating bonuses for basic abilities and positive behaviors. Average response time (D), number of consecutive requests for help (H), and number of consecutive incorrect answers (E) are negative factors, and therefore the summation sign is negative, indicating deductions for excessive time consumption, frequent requests for help, and consecutive incorrect answers. Through normalization and weighted summation, multi-dimensional heterogeneous parameters are integrated into a single quantitative indicator, facilitating subsequent level mapping.
[0040] After obtaining the comprehensive cognitive score, the score is mapped to five cognitive levels (L0-L4) according to a preset level interval table, resulting in the corresponding user cognitive level. This preset level interval table defines the mapping relationship between the comprehensive cognitive score and the user's cognitive levels L0 to L4, where L0-L1 corresponds to scenarios with low comprehension or high need for assistance, L2 corresponds to the standard interaction level, and L3-L4 corresponds to scenarios with high comprehension or low need for assistance. This mapping relationship can be calibrated according to actual educational scenarios, and the user's cognitive level can be recalculated based on the latest interaction behavior data after each round of interaction to achieve dynamic updates. For example, if a child answers correctly with a high success rate, responds quickly, and does not ask for help in multiple consecutive rounds, their comprehensive cognitive score will rise, potentially upgrading from L2 to L3; conversely, if they answer incorrectly repeatedly or frequently ask for help, their score will decrease, potentially downgrading to L1. This dynamic update mechanism allows for real-time adaptation to changes in the user's learning status, better assessment of the user's cognitive level, and more accurate and appropriate interactive feedback.
[0041] In the above embodiments, the present invention discloses a puzzle interaction control method for an intelligent reading pen. When a puzzle piece recognition event is triggered, the method acquires the identification information of the target puzzle piece; acquires the theme resource data corresponding to the target puzzle piece based on the identification information; acquires the user profile data and puzzle context data of the current user; determines the current interaction mode based on the received user input data and the puzzle context data; matches a corresponding interaction template according to the current interaction mode; injects the user profile data and puzzle context data as constraint parameters into the interaction template; and generates a corresponding initial response based on the theme resource data. The initial response is then adjusted for cognitive matching based on the user's cognitive level in the user profile data, generating and outputting a final response. By integrating puzzle piece recognition with user profiles and context for interaction mode switching, and dynamically adjusting the response content based on cognitive level, the same puzzle piece can output differentiated content depending on different situations, achieving dynamic interaction of the puzzle reading pen and improving the interactive adaptability and diversity of the puzzle teaching aid.
[0042] In one embodiment, step S203 includes: Perform action parsing and / or semantic parsing on the received user input data to obtain the corresponding pen tip action signal and / or semantic type; Determine the user's interaction intent based on the pen tip action signal and / or semantic type; Extract the historical recognition count and the most recent interaction record of the target puzzle piece from the puzzle context data, and confirm the current interaction scenario based on the historical recognition count and the most recent interaction record; The current interaction mode is determined based on the user's interaction intent and the current interaction scenario.
[0043] In this embodiment, the received user input data undergoes action analysis and / or semantic analysis. Action analysis involves thresholding and pattern recognition of the raw physical signals collected by the pen tip sensor, including pressure value change curves, press duration, and time intervals between presses, outputting standardized pen tip action signals. These signals include single clicks, double clicks, long presses, continuous presses, pause durations, and re-trigger intervals. Semantic analysis involves speech recognition and natural language understanding processing of the voice data collected by the microphone to determine the semantic type of the voice data. This semantic type includes whether it contains interrogative words, locative words, pronouns, requests for help, and confirmation words. This dual analysis of action and semantics captures the user's subjective expression and pre-set pen tip trigger actions, providing multi-dimensional input for subsequent intent determination.
[0044] The user's interaction intent is determined by comprehensively analyzing the pen-tip action signals and / or semantic types obtained from the analysis. The user's interaction intent refers to the category of goals the user currently hopes to achieve through interaction, including acquiring a story, seeking knowledge, requesting help, confirming right or wrong, and engaging in casual conversation, etc. Specifically, the action type and semantic type can be combined and mapped to the interaction intent using preset intent mapping rules. That is, an intent mapping table is maintained in advance, containing the corresponding user interaction intent for different pen-tip action signals and semantic types. For example, when a long press, two consecutive long presses, or a voice message containing help-seeking phrases like "Where to put it?", "Help me", or "Which piece is it in?" is detected, it is mapped to a "request for help intent". When a voice message contains confirmation phrases like "Is it?", "Is it right?", or "Can I put it here?", it is mapped to a "confirmation intent". When a question word is detected and the semantic match between the question and the current block tag exceeds a preset threshold, it is mapped to a "knowledge-seeking intent". When a short press is detected without a voice question, it is mapped to a "story-seeking intent". When the same puzzle piece is repeatedly pressed, it is mapped to a "repeated attention intent". When a question is asked across blocks and involves both the current and previous blocks, it is mapped to a "cross-block association intent", and so on. By jointly determining actions and semantics, the current user interaction intent is accurately obtained, providing a basis for judging interaction patterns.
[0045] Based on confirming the user's interaction intent, the current interaction scenario is identified according to the puzzle context, i.e., the temporal state of the target puzzle piece in the current session. This current interaction scenario includes initial recognition, repeated recognition, and cross-piece switching scenarios, etc. Specifically, the historical recognition count and the most recent interaction record of the target puzzle piece are extracted from the puzzle context data. For example, all historical recognition records within the current context window are read. The specific scope of the context window can be preset to a default value, such as the last 10 recognition events, the last 5 minutes, etc., or it can be dynamically adjusted according to the user interaction frequency. These historical recognition records are sorted in descending order by recognition timestamp. Then, all recognition records related to the target puzzle piece are filtered out based on the identification information of the target puzzle piece, and the recognition count is counted to obtain the historical recognition count of the target puzzle piece. At the same time, the record with the latest timestamp in the filtering results is extracted as the most recent interaction record, and information such as the interaction status, feedback content, and recognition time in the record is obtained. If there is no historical recognition record for the puzzle piece in the context window, the historical recognition count is determined to be 0, and the most recent interaction record is empty. By extracting the historical recognition count and the most recent interaction record of the target puzzle piece, historical evidence is provided for the decision-making of the current interaction mode, thus achieving contextual coherence.
[0046] Based on the historical recognition count of the target puzzle piece and the most recent interaction record, determine whether the current interaction scenario is a first recognition scenario, a repeated recognition scenario, or a cross-piece switching scenario, so as to output targeted puzzle interaction feedback. The specific criteria for classification are the number of historical recognitions, the ID of the most recently recognized puzzle piece, and the recognition time interval, ensuring the accuracy and rationality of the scenario classification. For example, the first recognition scenario refers to a scenario where the historical recognition count of the target puzzle piece is 0, or although there is a historical recognition record, the interval between the most recently recognized time and the current recognition time exceeds a preset threshold Tw1 (e.g., 5 minutes), and the historical context has expired. The repeated recognition scenario refers to a scenario where the unique block ID of the target puzzle piece is the same as the ID of the most recently recognized puzzle piece, and the time interval between the two recognitions is less than a preset threshold Tw2 (e.g., 30 seconds), or the number of recognitions of the block in the current context window is ≥2. The cross-block switching scenario refers to a scenario where the ID of the currently recognized target puzzle piece is different from the ID of the most recently recognized puzzle piece, and the time interval between the two recognitions is less than a preset threshold Tw3 (e.g., 1 minute), which is considered to be a block switching in the same continuous interaction link. Tw1, Tw2, and Tw3 can all be adjusted according to actual interaction needs, and this embodiment does not limit them.
[0047] Then, using a pre-defined decision rule engine, the current interaction mode is retrieved based on the user's interaction intent and the current interaction scenario. Specific decision rules include: If the user's interaction intent is "request for help," regardless of the current interaction scenario, the system enters the location assistance mode. If the user's interaction intent is "confirmation of correctness" and the system holds the basis for judging the current puzzle position, the system enters the error correction confirmation mode. In the first recognition scenario, if there is no clear interaction intent, the system corresponds to the first explanation mode; if there is a knowledge query intent, the system corresponds to the knowledge question and answer mode; if there is a request for help, the system corresponds to the location assistance mode, etc. In the repeated recognition scenario, if there is no clear interaction intent, the system corresponds to the repeated recognition review mode; if there is a request for help, the system corresponds to the location assistance mode. The system also switches between the supplementary explanation mode, review question mode, and encouragement feedback mode based on the number of repetitions, time window, and historical output coverage. In the cross-block switching scenario, if there is a cross-block association intent, the system corresponds to the cross-block association mode; if there is no clear interaction intent, the system corresponds to the first explanation mode (basic explanation of the new block), etc. The specific decision rules can be flexibly adjusted according to actual needs, and this embodiment does not limit them.
[0048] Preferably, when switching modes, the location assistance mode and error correction confirmation mode have higher priority than the normal narration and free chat modes. Furthermore, when the semantic recognition confidence level is lower than a preset lower limit, the interaction mode is not switched directly; instead, a clarifying statement is output or the user is asked to press the relevant puzzle piece again to ensure the reliability of the interaction mode's decision-making.
[0049] This embodiment identifies user intent and divides different interaction scenarios. Based on the dual constraints of user intent and interaction scenario, it determines the interaction mode, ensuring the reliability of the interaction mode determination. This allows the current interaction mode to be highly adapted to user needs and interaction scenarios, improving the adaptability of the puzzle reading interaction.
[0050] In one embodiment, step S204 includes: According to the current interaction mode, a corresponding interaction template is matched from a preset template set; the preset template set includes at least one response template under different interaction modes; The user profile data and puzzle context data are injected as constraint parameters into the interaction template to obtain a generation task with corresponding constraints. Based on the complexity of the generated task and the local resource hit data, a routing selection is made between local processing and cloud-enhanced processing to execute the generated task in conjunction with the topic resource data, thereby obtaining the initial response.
[0051] In this embodiment, the preset template set is a pre-built structured template library, categorized and stored according to interaction modes. Each interaction mode contains at least one response template. For example, the storytelling mode may include response templates such as character introduction templates, scene description templates, and association extension templates; the knowledge question-and-answer mode may include factual answer templates, causal explanation templates, and heuristic guidance templates. Each response template is a content framework with constraint slots and generation rules, enabling the generation of different response content based on personalized constraint parameters. Specifically, the corresponding category is retrieved from the preset template set according to the current interaction mode, and further filtered according to the puzzle theme and user age group to match the most suitable interaction template.
[0052] User profile data and puzzle context data are injected as constraint parameters into the currently matched interaction template to obtain a generated task with corresponding constraints. Specifically, parameters such as cognitive level, help need indicators, and parent mode parameters from the user profile data, and the current piece ID, most recent piece ID, current progress, candidate empty area, most recent prompt level, and most recent answer content from the puzzle context data, can be filled into the reserved slots of the interaction template or passed to the generation engine as external constraints. For example, after the "most recent piece ID" from the puzzle context is injected into the template, the generated task will include constraints such as "if the current answer involves the content of the previous piece, the pronoun 'it' must be used to refer to it and to maintain contextual coherence." By injecting constraint parameters, a general response template is transformed into a personalized generation task for the current user and current progress, thereby improving the adaptability between the response content and user needs.
[0053] After obtaining a generation task with constraints, the system routes between local processing and cloud-based enhanced processing based on task complexity and local resource availability, selecting different task execution entities. Specifically, the complexity of the generation task is automatically assessed based on factors such as whether the question involves cross-block reasoning, whether open-ended generation is required, whether the semantic length exceeds a threshold, and whether dynamic rewriting based on user profiles is needed. It also checks whether there are pre-set answers, rule outputs, or cached fixed questions and answers matching the current generation task in the local cache. If the complexity of the generation task is lower than the preset complexity, such as initial storytelling or standard knowledge point broadcasting, and relevant answers are already cached locally, it is routed to the local rule engine for local processing to obtain the corresponding initial response. If the complexity of the generation task is higher than the preset complexity, such as questions involving open-ended questions, cross-block related questions and answers, multi-turn contextual reasoning, questions with non-standard user expressions requiring semantic correction, or answers requiring dynamic rewriting based on user profiles, it is routed to the cloud server for cloud-based enhanced processing to obtain the initial response. In addition, a network degradation mechanism is set up. For example, if the processing in the cloud times out or the network is interrupted, it will automatically fall back to the local backup template to generate a reply, ensuring that the interaction is not interrupted.
[0054] In one embodiment, step S205 includes: Based on the user cognitive level in the user profile data, determine the corresponding difficulty index, help index, and follow-up question index; The knowledge density and vocabulary difficulty in the initial response are adjusted according to the difficulty index; the granularity of the location prompts in the initial response is adjusted according to the help index; and whether to add heuristic follow-up questions in the initial response is controlled according to the follow-up question index. The response content adjusted by cognitive matching is output as the final response.
[0055] In this embodiment, when adjusting the initial response based on cognitive matching, three decision control variables are first determined based on the user's cognitive level in the user profile data: a difficulty index Df, a help index Hf, and a follow-up question index Qf. These variables are used to quantify and control different dimensions of the output content. The difficulty index Df is a discrete or continuous value used to characterize the knowledge density and vocabulary complexity that the output content should achieve; the help index Hf is a discrete value (such as multi-level coding) used to characterize the level of detail in location prompts or error correction guidance; and the follow-up question index Qf is a Boolean or discrete value used to characterize whether a heuristic follow-up question is added after the current response.
[0056] When acquiring the three-dimensional metrics, a mapping table from user cognitive level to the metrics can be pre-defined. When the user's cognitive level is L0-L1, the difficulty metric Df is low, the help metric Hf is high, and the follow-up question metric Qf is off. When the user's cognitive level is L2, the difficulty metric Df is medium, the help metric Hf is medium, and the follow-up question metric Qf is optional. When the user's cognitive level is L3-L4, the difficulty metric Df is high, the help metric Hf is low, and the follow-up question metric Qf is always on. These three metrics can control word replacement, sentence length control, prompt level selection, and follow-up question on / off during cognitive matching adjustments, thus transforming the abstract cognitive level into directly executable control signals.
[0057] Specifically, this refers to adjusting the knowledge density and vocabulary difficulty in the initial response based on the difficulty index Df. This includes replacing words with synonyms, splitting or merging sentences, and increasing or decreasing the number of knowledge points. For example, when the difficulty index Df is low, high-level abstract words are replaced with concrete words (such as replacing "predation" with "catching small animals to eat," and "carnivorous animal" with "animal that eats meat"), complex sentences are broken down into short sentences, and multiple knowledge points are simplified into a single knowledge point statement. When the difficulty index Df is high, abstract words are retained or newly introduced, and complex sentences are used to express causal relationships and connect multiple knowledge points.
[0058] The granularity of location prompts in the initial response is adjusted based on the help index Hf. In this embodiment, when responding to a user's location request, the exact location is not given directly. Instead, multiple prompt levels are preset, including global area prompts, local semantic prompts, neighboring block prompts, shape matching prompts, and direct location prompts. The granularity of location prompts, i.e., the target prompt level, is adjusted based on the different users' help index Hf, thereby adjusting the level of detail in the location prompts. For example, when the help index Hf is high, the granularity of location prompts is reduced to increase the level of detail, such as changing from the local semantic "place it in the grassland area" to the neighboring block prompt "place it next to the zebra," or to the shape matching prompt "find a piece with an uneven edge," etc. When the help index Hf is low, heuristic and delayed help is prioritized, such as "you can first look at the color of this piece, and then find where you need this color," etc., thus providing personalized prompt granularity for users with different cognitive levels and at different stages of seeking help.
[0059] The Qf indicator controls whether to add heuristic follow-up questions to the initial response, i.e., questions that guide children's thinking. Specifically, when the Qf indicator is on, the initial response will end with questions such as "Would you like to try again?", "Who do you think it might be with?", "Do you know why?" etc. When the Qf indicator is off, the current answer will be completed without adding any additional follow-up questions.
[0060] After rewriting the content across multiple dimensions, the rewritten content undergoes compliance verification, such as checking for length limits, content safety, and age-inappropriate information. Once compliance verification is passed, it becomes the final response, output to the user in voice format. At this point, the final response is fully adapted to the user's current cognitive level and the current interaction scenario. By breaking down abstract cognitive levels into multiple quantifiable and adjustable control indicators, the initial response is rewritten to adapt to different learning paces, considering content difficulty, granularity of prompts, and dialogue continuity. This allows for personalized adjustments to the response content, better aligning with the learning rhythms of different children.
[0061] In one embodiment, determining the corresponding help indicator based on the user's cognitive level in the user profile data includes: Obtain the overall structure information of the entire jigsaw puzzle and identify the distribution information of the currently completed areas; Based on the puzzle piece characteristics of the target puzzle piece and the distribution information of the currently completed area, a corresponding puzzle state heterogeneous graph is constructed; The heterogeneous graph of the puzzle state is processed by a graph attention network to perform message passing and attention reasoning, thereby obtaining the matching attention score between the target puzzle piece and each missing position node, and the candidate region of the target puzzle piece is determined based on the matching attention score. Based on the user's cognitive level and the convergence degree of the candidate region, a corresponding target prompt level is selected from multiple preset prompt levels as the help indicator.
[0062] In this embodiment, when obtaining the help index, which represents the target prompt level representing the granularity of the prompt, it is obtained by combining the user's cognitive level and the current completion status of the puzzle. Specifically, the overall structure information of the entire puzzle is first obtained, and the distribution information of the currently completed areas is identified. The overall structure information is pre-established during the puzzle content creation stage and describes relevant information about the entire puzzle, used to store data such as the region division, object relationships, background structure, and puzzle piece position mapping of the entire puzzle. Based on this overall structure information, the static overall structure template is aligned with the dynamic completion status through one or more methods, such as the pressure sensor array of the puzzle base plate, the visual recognition module, or the position mapping table for real-time detection. This accurately identifies which positions have been correctly occupied, which positions are still empty, and the adjacency relationship between placed pieces and empty positions, thereby identifying the distribution information of the currently completed areas.
[0063] Then, based on the puzzle piece features of the target puzzle piece and the distribution information of the currently completed areas, a corresponding puzzle state heterogeneous graph is constructed. This puzzle state heterogeneous graph converts the actual completion state of the current puzzle into a graph structure data representation. Specifically, it abstracts puzzle pieces and empty positions in the physical space as graph nodes, and abstracts the multi-dimensional relationships between them as heterogeneous edges. The completed and empty areas are obtained based on the current distribution information of the completed areas. Each placed puzzle piece in the completed area, each empty position in the empty area, and the target puzzle piece are used as graph nodes. The node features of each graph node include semantic embedding vectors, color embedding vectors, edge encoding vectors, and direction encoding vectors. Placed puzzle piece nodes carry rich physical features, empty position nodes carry expected features, and the target puzzle piece node serves as a query anchor. These three types of nodes are interconnected through multi-dimensional heterogeneous edges.
[0064] When constructing heterogeneous edge relationships, corresponding edges are established based on the spatial adjacency, semantic continuity, and color compatibility relationships between graph nodes. Spatial adjacency edges are established based on the physical adjacency of positional units. For example, if the positional units corresponding to two nodes share an edge (up, down, left, right) in the grid, a strong spatial adjacency edge is established with a weight of 1.0; if they only share a vertex (diagonally), a weak spatial adjacency edge is established with a weight of 0.5. Semantic continuity edges are established based on hierarchical matching of semantic labels. If the semantic labels of two nodes share the same parent concept (e.g., "bird / wings" and "bird / body" share "bird"), a semantic continuity edge is established with a weight of 0.8; if they belong to the same background region (e.g., "sky / clouds" and "sky / sun"), the weight is 0.6, etc. Color compatibility edges can be established based on the Euclidean distance in the Lab color space. If the Lab distance between the primary colors of two nodes is less than a preset threshold, such as 30, a color compatible edge is established, with higher weights for colors that are more similar. By constructing graph nodes and various types of edges between them, the multi-dimensional topological relationships of the current puzzle state are fully represented, providing a reliable basis for subsequent candidate region reasoning and determination.
[0065] Based on the constructed heterogeneous graph of the puzzle state, a graph attention network (GAT) is used for message passing and attention reasoning to obtain the matching attention scores between the target puzzle piece and each empty position node. Candidate regions for the target puzzle piece are then determined based on these matching attention scores. The GAT is a graph neural network based on an attention mechanism, capable of adaptive message passing on heterogeneous graphs. Specifically, it uses a multi-head attention mechanism, allowing each empty position node to dynamically calculate the contribution weights of different neighboring nodes based on the features of its already placed puzzle piece nodes. This aggregates a high-order feature representation containing contextual semantics. Finally, based on the matching attention scores between the target puzzle piece and each empty position node, the top-k regions (e.g., the top 5) are selected as candidate regions for the target puzzle piece. The message propagation and reasoning of the GAT effectively improves the accuracy and generalization ability of candidate region determination by outputting the matching probability of each empty position as the correct location of the target piece.
[0066] The target prompt level is selected from multiple preset prompt levels based on the user's cognitive level and the convergence of candidate regions. The convergence of candidate regions refers to the concentration of attention scores among the Top-K candidate regions, specifically the ratio of the probability of the Top-1 candidate to the number of candidate regions. For example, if the Top-1 probability is 0.85 and the number of candidates is 5, the convergence is 0.17. A high convergence indicates a high degree of certainty about the target location; for example, the Top-1 score is significantly higher than others. Conversely, a low convergence indicates the existence of multiple possible locations, i.e., a uniform score distribution. When selecting the target prompt level, the lower the user's cognitive level (L0-L1), the higher the need for assistance; the higher the user's cognitive level (L3-L4), the lower the need for assistance. The convergence of candidate regions also determines the certainty of the location; a higher convergence is more suitable for providing precise prompts, while a lower convergence is more suitable for providing rough prompts to avoid misleading the user. Therefore, the target prompt level is selected from the preset prompt levels by combining both factors. For example, when the user's cognitive level is L0-L1 (high need for help) and the candidate region convergence is high, direct location prompts or neighboring block prompts are selected; when the user's cognitive level is L3-L4 (low need for help), even if the convergence is high, coarse region prompts or heuristic delayed prompts (such as "You can look at the color of this area first") are preferred. Through joint decision-making based on cognitive level and convergence, a balance between personalized and precise prompt granularity is achieved.
[0067] In one embodiment, after step S205, the method further includes: Collect user interaction behavior data for the current round, and update the puzzle context data and user profile data based on the interaction behavior data for use in the next round of interaction.
[0068] In this embodiment, after each round of interaction, the puzzle context data and user profile data are updated based on the interaction behavior data of the current round, so that the context data and profile can continuously reflect the user's latest state. Specifically, after the final reply is output, the user's interaction behavior data is collected in real time, including whether the answer is correct, whether a request for help is made, response time, voice input content, type of pressing action, listening time of the final reply, and puzzle piece placement status, etc.
[0069] When updating the puzzle context data based on interaction behavior data, the process specifically involves marking the current puzzle piece as the most recently interacted piece and updating its historical recognition count, updating the current progress percentage, recording the current interaction mode and output content summary, and updating the distribution of completed areas, etc. When updating user profile data, the accuracy rate is updated using a sliding window, retaining the data from the most recent N rounds (e.g., N=10) and calculating the new accuracy rate; the number of consecutive requests for help, the number of consecutive incorrect answers, and the average response time are updated exponentially; and the frequency of user-preferred topics and weaknesses is statistically analyzed to establish topic preference distributions and weakness distributions. For example, if the "African savanna" topic has a high interaction frequency, its preference weight is increased; if multiple requests for help are made in the lower right corner area, it is marked as a weakness, etc.
[0070] The updated puzzle context data and user profile data are written to a local cache database as input data for the next round of interaction. Through real-time data updates after each round of interaction, the dynamic evolution of user profiles and continuous updates of puzzle context state are achieved, enabling the next round of interaction to make decisions based on the latest context and factors such as the user's current cognition and preferences, thereby better supporting the coherence of continuous multi-round dialogues.
[0071] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0072] To better understand the puzzle interaction control process of the smart reading pen provided by this invention, the following specific application examples illustrate the puzzle interaction control process: Example 1 When a user first presses the reading pen on the "Deer in the Forest" puzzle piece, triggering a puzzle piece contact event, the reading pen's built-in NFC module reads the identification information (e.g., FOREST_THEME_007) from the puzzle piece's electronic tag. Based on this identification information, it retrieves corresponding theme resource data from the local content library, including the deer character's background story, the forest scene's semantic tags, and the piece's expected coordinates within the overall image. Simultaneously, it acquires the user's profile data (e.g., cognitive level L2, age 5, intermediate language ability) and puzzle context data (current progress 0%, this piece is being recognized for the first time). Since the user's input data is parsed as a single short press without any voice prompts, based on the rule of "first recognition without accompanying questioning," the current interaction mode is determined to be storytelling mode.
[0073] In storytelling mode, a storytelling template is matched, and user cognitive level L2, the semantic label of the current block "deer", and the scene "forest" are injected into the template as constraint parameters. Since this task is a standard structured storytelling and relevant materials are cached locally, it is routed to the local processing unit, which generates an initial response based on the topic resource data: "The deer lives deep in the forest. It has brown fur and long ears, and it likes to drink water by the stream." In addition, based on user cognitive level L2, the difficulty index is determined to be "medium" and the follow-up question index is set to "on". A heuristic follow-up question is added at the end: "Do you know why the deer's ears are so long?" The final response is obtained and output through speech synthesis, thus transforming the traditional static puzzle block recognition and fixed audio playback into a dynamic and interactive age-appropriate storytelling.
[0074] Example 2 After the story was narrated in Example 1, the user asked, "Why is it in the forest?" At this point, the voice input was collected and semantically analyzed, identifying the question word "why" and the semantic type as "seeking knowledge". Simultaneously, the semantic tag "deer" of the current block was extracted from the puzzle context data and matched with the question semantics to determine that the current interaction mode is a knowledge question-and-answer mode.
[0075] At this point, a knowledge-based question-and-answer template is matched, injecting the current puzzle theme "forest ecosystem," the user's cognitive level L2, and a summary of the previous presentation as constraint parameters. This generation task involves ecosystem knowledge points, and since the local cache did not contain complete causal chain materials, it is routed to the cloud for augmentation processing. The cloud model combines the forest ecosystem knowledge points from the theme resource data to generate an initial response, which is then matched and adjusted based on the user's cognitive level before outputting the final response.
[0076] Example 3 If a user presses and holds the pen tip for more than 1.5 seconds while completing a puzzle, and asks, "Where should this piece be placed?" the action is identified as a "long press" and the semantics are identified as a "help request" or "location word". Therefore, the user's interaction intent is determined to be a "request for help" and the current interaction mode is determined to be the location help mode with the highest priority.
[0077] At this point, the overall structure information of the entire puzzle and the distribution information of the currently completed areas are obtained. For example, the upper half of the forest area is completed, with 3 gaps remaining. A heterogeneous graph of the puzzle state is constructed using the puzzle piece features (semantic labels, edge shape encoding, etc.) of the target puzzle piece "deer" and the distribution information of the completed areas. Nodes include placed pieces and gap positions, and edges cover spatial adjacency, semantic association, and shape matching relationships. Message passing and attention inference are performed through a graph attention network to calculate the matching attention score between the target puzzle piece and each gap position node. The inference results show that the attention score of the gap position (3,3) is 0.82, significantly higher than other candidate positions, indicating a high degree of convergence of the candidate region. Considering that the user's cognitive level is L1 (high help demand), "neighboring piece prompts" is selected as the target prompt level from the preset prompt levels, and the help index is encoded as "neighboring piece prompt level".
[0078] Therefore, when generating the response content, after matching the location prompt template and injecting constraint parameters to generate the initial response, the granularity of the prompts, sentence length, and vocabulary are adjusted based on the cognitive level. This ensures that the output of adjacent block prompts is achieved, and complex sentences are broken down into short sentences while retaining concrete vocabulary. The final response is: "The deer block should be placed in the middle area of the forest, near the large tree block you just placed. That block has a semi-circular notch on its edge, which fits perfectly with the convex edge of the deer block." This achieves layered prompts based on cognitive level, balancing puzzle completion efficiency with educational guidance.
[0079] Example 4 For the same puzzle piece, "Deer in the Forest," differentiated content output is achieved for users with different cognitive levels through dynamic user profiling and cognitive matching adjustment mechanisms.
[0080] For young children (User A, age 3, cognitive level L0), their basic profile parameters (age 3, low language ability) and interaction behavior parameters (2 consecutive requests for help, 30% accuracy rate) are obtained. The difficulty index is set to "low", the help index to "high prompt level", and the follow-up question index to "closed". Therefore, when adjusting cognitive matching after generating the initial response, vocabulary is replaced with concrete expressions, complex sentences are broken down into short sentences, and onomatopoeic encouragement is added.
[0081] For older children (User B, age 7, cognitive level L4), based on their high accuracy rate (95%), low frequency of seeking help, and fast response time, the difficulty index was set to "high," the help index to "low level of prompts," and the follow-up question index to "frequently enabled." Therefore, when generating the response content and adjusting for cognitive matching, abstract vocabulary ("deer species," "forest ecosystem," etc.) was retained, and complex sentences were used to connect the knowledge points, with an open-ended question appended at the end: "If the forest is cut down, what impact will it have on the survival of deer?" Example 5 If a user presses the "Deer in the Forest" puzzle piece for the first time using the reading pen in an environment without network access (such as during a long journey), and the network quality is detected to be below the threshold N1, according to the local priority strategy, the local processing unit directly retrieves the theme resource data from the local content library and executes the storytelling template generation task, outputting the cached story content: "The deer lives deep in the forest..." Meanwhile, user interaction data (including voice questions like "What does it eat?") and puzzle context data (current progress, set of placed pieces, etc.) are recorded in the offline session and temporarily stored in a local cache database. Since the network is unavailable at this point, tasks involving generating open-ended questions cannot be routed to the cloud. Therefore, the system reverts to the local basic answer: "The deer eats grass and leaves," and outputs the message: "This is an interesting question; I'll tell you in detail when the network is back up." Once the device regains network connectivity, the system automatically synchronizes the temporarily stored interaction data and puzzle context data from the offline period to the cloud server. The cloud model then supplements this data to generate richer question-and-answer responses (e.g., "Deer are herbivores, mainly feeding on tender grass, leaves, and bark, with leaves making up more than 60% of their diet..."), and pushes these responses to the reading pen terminal. Simultaneously, the cloud updates the puzzle progress and user profile data, enabling progress synchronization across multiple devices. This ensures the availability of basic interactions in offline scenarios and automatically completes enhanced content supplementation and data synchronization after network recovery.
[0082] Further reference Figure 3 As a response to the above Figure 2 The present invention provides an embodiment of a puzzle interaction control device for an intelligent reading pen, which, according to the method shown, provides a solution for puzzle interaction control. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0083] like Figure 3 As shown, the puzzle interaction control device 30 of the smart reading pen described in this embodiment includes: The jigsaw puzzle recognition module 301 is used to obtain the identification information of the target jigsaw puzzle piece when the jigsaw puzzle piece recognition event is triggered; Resource acquisition module 302 is used to acquire theme resource data corresponding to the target puzzle piece based on the identification information; The interaction mode switching module 303 is used to obtain the current user's user profile data and puzzle context data, and determine the current interaction mode based on the received user input data and the puzzle context data. The response generation module 304 is used to match the corresponding interaction template according to the current interaction mode, inject the user profile data and puzzle context data as constraint parameters into the interaction template, and generate the corresponding initial response in combination with the topic resource data; The cognitive adjustment module 305 is used to adjust the initial response based on the user's cognitive level in the user profile data, and generate and output the final response.
[0084] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the puzzle interaction control execution process of the smart reading pen. For the specific implementation of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0085] In one embodiment, the interaction mode switching module 303 includes: The parsing unit is used to perform action parsing and / or semantic parsing on the received user input data to obtain the corresponding pen tip action signal and / or semantic type. An intent determination unit is used to determine the user's interaction intent based on the pen tip action signal and / or semantic type; The scene determination unit is used to extract the historical recognition count and the most recent interaction record of the target puzzle piece from the puzzle context data, and to confirm the current interaction scene based on the historical recognition count and the most recent interaction record; The interaction mode switching unit is used to determine the current interaction mode based on the user's interaction intent and the current interaction scenario.
[0086] In one embodiment, the response generation module 304 includes: A template matching unit is used to match a corresponding interaction template from a preset template set according to the current interaction mode; the preset template set includes at least one response template under different interaction modes; The task generation unit is used to inject the user profile data and puzzle context data as constraint parameters into the interaction template to obtain a generated task with corresponding constraints. The response generation unit is used to select between local processing and cloud-enhanced processing based on the complexity of the generation task and the local resource hit data, so as to execute the generation task in combination with the topic resource data and obtain the initial response.
[0087] In one embodiment, the device 30 further includes: The parameter acquisition module is used to acquire basic profile parameters and interaction behavior parameters; The cognitive scoring module is used to normalize the basic profile parameters and interaction behavior parameters, and perform weighted summation of the basic profile parameters and interaction behavior parameters according to preset weight configuration to obtain a comprehensive cognitive score. The cognitive mapping module is used to map the comprehensive cognitive score to the corresponding user cognitive level according to a preset level range table.
[0088] In one embodiment, the cognitive adjustment module 305 includes: The indicator extraction unit determines the corresponding difficulty indicator, help indicator, and follow-up question indicator based on the user's cognitive level in the user profile data. The content adjustment unit is used to adjust the knowledge density and vocabulary difficulty in the initial response according to the difficulty index, adjust the granularity of the location prompts in the initial response according to the help index, and control whether to add heuristic follow-up questions in the initial response according to the follow-up question index. The response output unit is used to output the response content adjusted by cognitive matching as the final response.
[0089] In one embodiment, the indicator extraction unit includes: The progress recognition unit is used to obtain the overall structure information of the entire puzzle and identify the distribution information of the currently completed areas; The graph construction unit is used to construct a corresponding puzzle state heterogeneous graph based on the puzzle piece characteristics of the target puzzle piece and the distribution information of the currently completed area; The candidate unit is used to perform message passing and attention reasoning on the heterogeneous graph of the puzzle state through a graph attention network, obtain the matching attention score between the target puzzle piece and each empty position node, and determine the candidate region of the target puzzle piece based on the matching attention score; The help indicator determination unit is used to select a corresponding target prompt level from multiple preset prompt levels as the help indicator based on the user's cognitive level and the convergence degree of the candidate region.
[0090] In one embodiment, the device 30 further includes: The user profile update module is used to collect user interaction behavior data in the current round, and update the puzzle context data and user profile data according to the interaction behavior data for the next round of interaction.
[0091] In the above embodiments, the present invention discloses a puzzle interaction control device for an intelligent reading pen. When a puzzle piece recognition event is triggered, the device acquires the identification information of the target puzzle piece; obtains the theme resource data corresponding to the target puzzle piece based on the identification information; acquires the user profile data and puzzle context data of the current user; determines the current interaction mode based on the received user input data and the puzzle context data; matches a corresponding interaction template according to the current interaction mode; injects the user profile data and puzzle context data as constraint parameters into the interaction template; and generates a corresponding initial response based on the theme resource data. The initial response is then adjusted for cognitive matching based on the user's cognitive level in the user profile data, generating and outputting a final response. By integrating puzzle piece recognition with user profiles and context for interaction mode switching, and dynamically adjusting the response content based on cognitive level, the same puzzle piece can output differentiated content depending on different situations, achieving dynamic interaction of the puzzle reading pen and improving the interactive adaptability and diversity of the puzzle teaching aid.
[0092] Specific limitations regarding the puzzle interaction control device of the smart reading pen can be found in the limitations of the puzzle interaction control method for the smart reading pen mentioned above, and will not be repeated here. Each module in the aforementioned puzzle interaction control device of the smart reading pen can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0094] Another embodiment of the present invention provides a computer device, such as... Figure 4 As shown, the computer device 40 includes: One or more processors 401 and memory 402, Figure 4 The following section uses a processor 401 as an example. The processor 401 and the memory 402 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.
[0095] The processor 401 is used to perform various control logics of the computer device 40. It can be any conventional processor, microprocessor, state machine, general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0096] The memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the puzzle interaction control method of the intelligent reading pen in the embodiments of the present invention. The processor 401 executes various functional applications and data processing of the computer device 40 by running the non-volatile software programs, instructions, and units stored in the memory 402, thereby realizing the puzzle interaction control method of the intelligent reading pen in the above method embodiments.
[0097] Another embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, perform the steps of the puzzle interaction control method of the smart reading pen in any of the above method embodiments.
[0098] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0099] Based on the above description of the embodiments, those skilled in the art will understand that the methods described in the embodiments can be implemented using software plus necessary general-purpose hardware platforms. Of course, they can also be implemented using hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0100] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0101] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0102] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for controlling the puzzle interaction of an intelligent reading pen, characterized in that, include: When the puzzle piece recognition event is triggered, obtain the identification information of the target puzzle piece; Obtain the theme resource data corresponding to the target puzzle piece based on the identification information; Obtain the current user's user profile data and puzzle context data, and determine the current interaction mode based on the received user input data and the puzzle context data; Match the corresponding interaction template according to the current interaction mode, inject the user profile data and puzzle context data as constraint parameters into the interaction template, and generate the corresponding initial response in combination with the topic resource data; Based on the user's cognitive level in the user profile data, the initial response is adjusted for cognitive matching, and the final response is generated and output. The step of adjusting the initial response based on the user's cognitive level in the user profile data, and generating and outputting the final response, includes: Based on the user cognitive level in the user profile data, determine the corresponding difficulty index, help index, and follow-up question index; The knowledge density and vocabulary difficulty in the initial response are adjusted according to the difficulty index; the granularity of the location prompts in the initial response is adjusted according to the help index; and whether to add heuristic follow-up questions in the initial response is controlled according to the follow-up question index. The response content adjusted based on cognitive matching is output as the final response; The step of determining corresponding help indicators based on the user's cognitive level in the user profile data includes: Obtain the overall structure information of the entire jigsaw puzzle and identify the distribution information of the currently completed areas; Based on the puzzle piece characteristics of the target puzzle piece and the distribution information of the currently completed area, a corresponding puzzle state heterogeneous graph is constructed; The heterogeneous graph of the puzzle state is processed by a graph attention network to perform message passing and attention reasoning, thereby obtaining the matching attention score between the target puzzle piece and each missing position node, and the candidate region of the target puzzle piece is determined based on the matching attention score. Based on the user's cognitive level and the convergence degree of the candidate region, a corresponding target prompt level is selected from multiple preset prompt levels as the help indicator.
2. The puzzle interaction control method for the intelligent reading pen according to claim 1, characterized in that, Determining the current interaction mode based on the received user input data and the puzzle context data includes: Perform action parsing and / or semantic parsing on the received user input data to obtain the corresponding pen tip action signal and / or semantic type; Determine the user's interaction intent based on the pen tip action signal and / or semantic type; Extract the historical recognition count and the most recent interaction record of the target puzzle piece from the puzzle context data, and confirm the current interaction scenario based on the historical recognition count and the most recent interaction record; The current interaction mode is determined based on the user's interaction intent and the current interaction scenario.
3. The puzzle interaction control method for the intelligent reading pen according to claim 1, characterized in that, The step of matching a corresponding interaction template based on the current interaction mode, injecting the user profile data and puzzle context data as constraint parameters into the interaction template, and generating a corresponding initial response in conjunction with the topic resource data includes: According to the current interaction mode, a corresponding interaction template is matched from a preset template set; the preset template set includes at least one response template under different interaction modes; The user profile data and puzzle context data are injected as constraint parameters into the interaction template to obtain a generation task with corresponding constraints. Based on the complexity of the generated task and the local resource hit data, a routing selection is made between local processing and cloud-enhanced processing to execute the generated task in conjunction with the topic resource data, thereby obtaining the initial response.
4. The puzzle interaction control method for the intelligent reading pen according to claim 1, characterized in that, The user's cognitive level is obtained through the following steps: Obtain basic profile parameters and interaction behavior parameters; The basic profile parameters and interaction behavior parameters are normalized, and then weighted and summed according to preset weights to obtain a comprehensive cognitive score. The comprehensive cognitive score is mapped to the corresponding user cognitive level according to a preset level range table.
5. The puzzle interaction control method for the intelligent reading pen according to claim 1, characterized in that, After adjusting the initial response based on the user's cognitive level in the user profile data, generating and outputting the final response, the method further includes: Collect user interaction behavior data for the current round, and update the puzzle context data and user profile data based on the interaction behavior data for use in the next round of interaction.
6. A puzzle interaction control device for an intelligent reading pen, characterized in that, include: The puzzle recognition module is used to obtain the identification information of the target puzzle piece when the puzzle piece recognition event is triggered; The resource acquisition module is used to acquire the theme resource data corresponding to the target puzzle piece based on the identification information; The interaction mode switching module is used to obtain the current user's user profile data and puzzle context data, and determine the current interaction mode based on the received user input data and the puzzle context data; The response generation module is used to match the corresponding interaction template according to the current interaction mode, inject the user profile data and puzzle context data as constraint parameters into the interaction template, and generate the corresponding initial response in combination with the topic resource data; The cognitive adjustment module is used to perform cognitive matching adjustment on the initial response based on the user's cognitive level in the user profile data, and generate and output the final response. The cognitive adjustment module includes: The indicator extraction unit determines the corresponding difficulty indicator, help indicator, and follow-up question indicator based on the user's cognitive level in the user profile data. The content adjustment unit is used to adjust the knowledge density and vocabulary difficulty in the initial response according to the difficulty index, adjust the granularity of the location prompts in the initial response according to the help index, and control whether to add heuristic follow-up questions in the initial response according to the follow-up question index. The response output unit is used to output the response content adjusted by cognitive matching as the final response; The indicator extraction unit includes: The progress recognition unit is used to obtain the overall structure information of the entire puzzle and identify the distribution information of the currently completed areas; The graph construction unit is used to construct a corresponding puzzle state heterogeneous graph based on the puzzle piece characteristics of the target puzzle piece and the distribution information of the currently completed area; The candidate unit is used to perform message passing and attention reasoning on the heterogeneous graph of the puzzle state through a graph attention network, obtain the matching attention score between the target puzzle piece and each empty position node, and determine the candidate region of the target puzzle piece based on the matching attention score; The help indicator determination unit is used to select a corresponding target prompt level from multiple preset prompt levels as the help indicator based on the user's cognitive level and the convergence degree of the candidate region.
7. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the puzzle interactive control method of the smart reading pen according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the puzzle interactive control method of the smart reading pen according to any one of claims 1-5.
Citation Information
Patent Citations
Touch and talk pen interactive content generation method based on adaptive learning
CN119850377A
Devices, systems, and methods for eliciting behavioral change
US20250210168A1