Adaptive interaction method and device based on puzzle piece recognition, equipment and medium
By receiving puzzle piece recognition events and user input events, parsing puzzle piece information and user intent, and combining contextual history, determining the interaction state and outputting feedback, the problem of the single interaction logic of existing puzzle teaching aids is solved, and adaptive multi-interaction modes and intelligent puzzle interaction are realized.
Patent Information
- Application Number
- CN202610844418.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-06-11
AI Technical Summary
Existing recognition-based teaching aids struggle to achieve adaptive adjustment of multiple interaction modes when combined with physical puzzle objects, resulting in simplistic and fragmented interaction logic.
By receiving puzzle piece recognition events and user input events, the system parses the target puzzle piece information and user interaction intent, combines the context history to determine the current interaction state, outputs corresponding puzzle interaction feedback, and updates the context cache data.
It has improved the continuity and intelligence of puzzle interaction, overcoming the problems of traditional puzzle teaching aids that have fixed audio playback and no scene-adaptive interaction.
Smart Images

Figure CN122387325B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an adaptive interaction method, apparatus, device, and medium based on puzzle piece recognition. Background Technology
[0002] A reading pen is a reading and learning tool that uses optical image recognition and digital voice technology. It can produce sound by touching specific areas on a companion medium such as a book, puzzle, or doll. It is usually used as a recognition-based teaching tool for children's early education.
[0003] Currently, existing recognition-based teaching aids typically equate label recognition with playing fixed content, achieving only simple interaction of recognition and playback. Even if some products have voice interaction, they often operate independently of the physical reading object, resulting in a single and fragmented interaction logic for existing recognition-based puzzle teaching aids, making it difficult to achieve adaptive dynamic adjustment of multiple interaction modes. Summary of the Invention
[0004] The main objective of this invention is to provide an adaptive interaction method, apparatus, device, and medium based on puzzle piece recognition, aiming to solve the technical problem that existing recognition-based teaching aids are difficult to achieve adaptive adjustment of multiple interaction modes in combination with physical puzzle objects.
[0005] The technical solution of the present invention is as follows: The first aspect of this invention provides an adaptive interaction method based on puzzle piece recognition, comprising: Receive puzzle piece recognition events and user input events to be processed, wherein the user input events include pen tip action events and / or voice input events; The puzzle piece recognition event is parsed to obtain the target puzzle piece information, and the user input event is parsed to obtain the user's interaction intent; Obtain historical recognition records from the context window, and determine the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records; Output the corresponding puzzle interaction feedback based on the current interaction state, and update the context cache data.
[0006] A second aspect of the present invention provides an adaptive interactive device based on puzzle piece recognition, comprising: The event receiving module is used to receive puzzle piece recognition events and user input events to be processed, wherein the user input events include pen tip action events and / or voice input events; The event parsing module is used to parse the puzzle piece recognition event to obtain target puzzle piece information, and to parse the user input event to obtain user interaction intent; The interaction state control module is used to obtain historical recognition records in the context window and determine the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records. The interactive feedback module is used to output corresponding puzzle interaction feedback based on the current interaction state and update the context cache data.
[0007] A third aspect of the present invention provides a computer device including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the aforementioned adaptive interaction method based on puzzle piece recognition.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the above-described adaptive interaction method based on puzzle piece recognition.
[0009] Beneficial Effects: This invention discloses an adaptive interaction method, apparatus, device, and medium based on jigsaw puzzle piece recognition. Compared to existing technologies, this invention receives jigsaw puzzle piece recognition events and user input events, including pen action events and / or voice input events; parses the jigsaw puzzle piece recognition events to obtain target jigsaw puzzle piece information, and parses the user input events to obtain user interaction intent; obtains historical recognition records in the context window, and determines the current interaction state based on the target jigsaw puzzle piece information, user interaction intent, and historical recognition records; outputs corresponding jigsaw puzzle interaction feedback based on the current interaction state, and updates the context cache data. This invention, by recognizing jigsaw puzzle piece recognition events and user input events, and combining them with context history records to achieve intelligent switching of interaction states, overcomes the problem of traditional jigsaw puzzle teaching aids only playing fixed audio and lacking scene-adaptive interaction, effectively improving the continuity and intelligence level of jigsaw puzzle reading interaction. Attached Figure Description
[0010] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1This is a schematic diagram of an application environment for the adaptive interaction method based on puzzle piece recognition provided in an embodiment of the present invention. Figure 2 A flowchart of an adaptive interaction method based on puzzle piece recognition provided in an embodiment of the present invention; Figure 3 A schematic diagram of the functional modules of the adaptive interactive device based on puzzle piece recognition provided in an embodiment of the present invention; Figure 4 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.
[0013] The adaptive interaction method based on puzzle piece recognition provided in this invention can be applied to, for example... Figure 1 In the interactive scenario shown, where the smart reading pen is paired with the puzzle device, there are a terminal device 101, a smart reading pen 102, a network 103, and a server 104. The network 103 serves as a medium to provide a communication link between the terminal device 101, the smart reading pen 102, and the server 104. The network 103 can include various connection types, such as wired and / or wireless communication links (e.g., Bluetooth, Wi-Fi, NFC, etc.).
[0014] Users can use terminal device 101 and smart reading pen 102 to interact with server 104 via network 103 to receive or send messages, etc. A client supporting this interaction method can be installed on terminal device 101, allowing users to log in to view and edit the interaction settings of smart reading pen 102, view puzzle interaction records, etc. Terminal device 101 can be various electronic devices with a display screen and web browsing support, including but not limited to smartphones, tablets, and desktop computers.
[0015] The smart reading pen 102 is used in conjunction with the puzzle piece 105. When the user places the smart reading pen 102 on the puzzle piece 105, it can automatically recognize the identifiable labels on the surface of the puzzle piece, thereby triggering a recognition event and outputting corresponding interactive content in different interactive states. Simultaneously, the user can press the button on the smart reading pen 102 to input voice through its built-in microphone for intelligent voice dialogue, achieving accurate transmission of interactive needs. The smart reading pen 102 integrates motion sensing, near-field communication recognition, voice acquisition, and feedback output functions, enabling it to sense the user's touch, press, and other operations in real time and associate them as corresponding interactive events.
[0016] Server 104 can be a server providing various services, such as analyzing and processing data like puzzle information recognized by the smart reading pen 102 and user-inputted voice commands, generating AI interactive content, and feeding it back to the smart reading pen 102's backend server (this is just an example). Server 104 can parse and process the received puzzle recognition data and user interaction commands, generate adapted interactive feedback content, and then transmit it to the smart reading pen 102 via network 103, where the smart reading pen 102 provides feedback to the user in the form of voice, light, etc. Server 104 can be a cloud server, a distributed system server, or a server integrated with blockchain.
[0017] It should be understood that the number of terminal devices 101, smart reading pens 102, networks 103, and servers 104 mentioned above is merely illustrative. Depending on the implementation needs, there can be any number of terminal devices 101, smart reading pens 102, networks 103, and servers 104.
[0018] like Figure 2 As shown, the adaptive interaction method based on puzzle piece recognition provided in this embodiment of the invention specifically includes the following steps: S201. Receive puzzle piece recognition events and user input events to be processed, wherein the user input events include pen action events and / or voice input events.
[0019] In this embodiment, the puzzle piece recognition event refers to the event triggered when the reading pen touches the identification label (such as an RFID tag or visual identification code) on the surface of the puzzle piece, or it can be triggered by the camera on the reading pen capturing the image of the puzzle piece and recognizing the puzzle piece identifier. In addition to triggering the puzzle piece recognition event by bringing the reading pen close to the puzzle piece, the user can also trigger the pen tip action event by using a preset button on the reading pen, or input corresponding interactive voice through the microphone. These two methods can be triggered individually or in combination.
[0020] Specifically, the system can monitor the recognition and interaction signals of the reading pen in real time. When the reading pen's recognition end is close to the puzzle piece label and meets the recognition distance threshold (e.g., 0.5-2cm), a puzzle piece recognition event is triggered, and the system receives the recognition event. Simultaneously, it also collects pen tip movements in real time. When user actions such as pressing, double-clicking, or long-pressing are detected, a pen tip movement event is triggered. The system also monitors ambient voice in real time through the reading pen's built-in microphone. When a voice signal is detected and the volume exceeds a preset threshold (e.g., 60dB), a voice input event is triggered. The voice signal is converted into a digital signal and transmitted to the control module. Alternatively, the voice function can be activated only after a specific action is detected, such as when the user long-presses a button to activate the microphone and receive the user's voice input, thus avoiding accidental triggering of voice input events.
[0021] For example, when a child is assembling a "forest animals" puzzle, they can touch the "deer" puzzle piece with the reading pen, triggering a puzzle piece recognition event. At the same time, they can press and hold the pen tip, triggering a pen tip action event. At this time, the microphone is activated to receive the user's voice input, triggering a voice input event. By collecting the coordinated triggering of various interactive events, a reliable foundation is provided for clarifying the user's interactive intent, breaking through the limitation of traditional reading puzzles that can only passively recognize and play fixed audio.
[0022] S202. Analyze the puzzle piece recognition event to obtain the target puzzle piece information, and analyze the user input event to obtain the user's interaction intent.
[0023] In this embodiment, upon receiving a puzzle piece recognition event, the event can be parsed using a tag parsing algorithm (such as RFID read / write algorithm, visual recognition algorithm, etc.) to obtain target puzzle piece information, such as the unique block ID of the target puzzle piece, as well as the block name, its theme, related knowledge points, and information about adjacent puzzle pieces. For user input events, action parsing and voice parsing are used to obtain the corresponding action type, voice content, and semantic analysis results, forming a clear user interaction intent. This ensures that, based on accurately locating the target puzzle piece for the current interaction, the user's physical actions or voice requests are transformed into identifiable interaction intents, achieving a precise mapping from user actions to interaction intents.
[0024] For example, after receiving the recognition event of the "deer" puzzle piece, the block ID is parsed to be "001". At the same time, the user's voice input "What does it eat?" is received. The semantic analysis is used to identify the knowledge question and answer semantics, and then the user's interaction intent is determined to be "to query knowledge points about the diet of deer", so as to realize diverse interaction scenarios.
[0025] S203. Obtain the historical recognition records in the context window, and determine the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records.
[0026] In this embodiment, the context window is a preset time window or session window used to store recent interaction data, such as the last 10 recognition events, interaction records within the last 5 minutes, etc. All historical recognition records within the current context window are acquired, including past puzzle piece recognition sequences, the number of times each puzzle piece was recognized, the corresponding interaction state and feedback content for each recognition, etc., to provide a historical event sequence basis for judging the current interaction state transition and ensure the continuity of the interaction context. Based on the acquired historical recognition records, combined with the parsed target puzzle piece information and user interaction intent, the current interaction state is determined. Specifically, a state machine can be used to maintain multiple interaction states and switch between them according to preset rules. That is, by using the historical interaction data of the target puzzle piece and the current user interaction intent, the optimal state transition path is matched according to the corresponding state transition rules, thereby confirming the current interaction state. For example, if the historical recognition record in the context window shows that the "deer" puzzle piece (ID001) has not been recognized before, and the user's current interaction intent is "querying knowledge points", the current interaction state is determined to be "knowledge question and answer state" according to the state transition rule, and feedback is prepared to be output in combination with the knowledge question and answer needs; if the historical recognition record shows that the "deer" puzzle piece has been recognized twice, and the most recent interaction state is "first narration state", and the user's current interaction intent is "asking for help", then the current interaction state is determined to be "location help state". This achieves adaptive switching between different interaction states such as narration, question and answer, and location help based on factors such as puzzle piece, user action form, historical recognition record, and current questioning intent, while maintaining contextual coherence.
[0027] S204. Output the corresponding puzzle interaction feedback based on the current interaction state, and update the context cache data.
[0028] In this embodiment, the puzzle interaction feedback is targeted content output to the user based on the current interaction state, including voice explanations, text prompts, and light guidance, to accurately match the user's current interaction intent and puzzle progress, avoiding repetitive output and invalid feedback. Specifically, it can call preset feedback templates based on the current interaction state, such as a "basic knowledge point explanation template" for the first explanation state, a "question and answer template" for the knowledge Q&A state, and a "layered prompt template" for the location help state, etc., and then combine the target puzzle piece information and the user's interaction intent to generate personalized feedback content. Alternatively, a large language model can be used to generate output feedback on the target puzzle piece and the user's interaction intent in the current interaction state. After outputting the puzzle interaction feedback, the current interaction state, the output feedback content, the target puzzle piece information, the user's interaction intent, and other data are written to the context cache to overwrite old redundant data, ensuring that subsequent interactions can make decisions based on the latest context information.
[0029] For example, if the current interaction state is "first time telling the story" and the user's interaction intent is "to query knowledge points", then the question-and-answer template is invoked, and feedback content "This is a deer. It is a herbivore that likes to eat leaves and grass and lives in the forest" is generated by combining the target puzzle piece information of "deer". This content is then output through the speaker of the reading pen. At the same time, information such as "first time telling the story", "feedback content" and "deer ID001, recognition count 1" are updated to the context cache, waiting for the next round of event triggering.
[0030] In one embodiment, step S202 includes: The puzzle piece recognition event is parsed to extract the unique block ID of the target puzzle piece; The pen tip action events are parsed to identify the pen tip action type and encode it into corresponding interactive control commands; Perform semantic parsing on voice input events to identify help-seeking semantics, confirmation semantics, knowledge-based question-and-answer semantics, and / or cross-block association semantics in the current input voice; The user's interaction intent is determined based on the interaction control instructions and / or semantic parsing results.
[0031] In this embodiment, upon receiving a puzzle piece recognition event, the recognition signal transmitted by the reading pen is decoded through tag parsing to extract the unique block ID stored in the tag. If parsing fails, such as no block ID is recognized or the parsed block ID is invalid, a prompt message such as "Please realign the puzzle piece" can be output, and the system returns to the listening state, waiting for the user to trigger the recognition event again. If parsing fails several times in a row, a prompt message "The puzzle piece may be damaged, please check" is output to improve the user experience.
[0032] When analyzing pen tip action events, the type of pen tip action can be identified based on the duration and number of touches of the pressing signal. For example, a single click refers to a touch time of 0.1-1 seconds and one touch; a double click refers to two touches within 1 second, with each touch lasting no more than 0.5 seconds; a long press refers to a touch time exceeding 2 seconds and one touch, and so on. The specific time and number of touches can be flexibly set. After identification, different pen tip action types are encoded into preset interactive control commands. For example, a single click is encoded as "CMD1" (e.g., corresponding to "play command"), a double click is encoded as "CMD2" (e.g., corresponding to "explain in another way"), and a long press is encoded as "CMD3" (e.g., corresponding to "help command"). This serves as the basis for determining the user's interaction intent and facilitates quick identification and execution of the corresponding interactive logic. It is understood that pen tip action types can be expanded according to actual needs, such as a triple click corresponding to a "repeat playback command" encoded as "CMD4", etc. This embodiment does not limit this.
[0033] For voice input events, semantic analysis is performed to accurately capture user interaction needs by recognizing key semantics in the speech. Specifically, "help" semantics correspond to users needing assistance with puzzle piece placement or assembly methods; "confirmation" semantics correspond to users needing to confirm whether puzzle pieces are placed correctly or whether they understand the concepts; "knowledge question" semantics correspond to users needing to query knowledge points related to puzzle pieces; and "cross-piece association" semantics correspond to users needing to inquire about the relationships between different puzzle pieces.
[0034] Specifically, user voice input can be fed into a pre-trained voice model. This model is pre-trained on a dataset containing commonly used children's voices and puzzle-related voices, and can quickly identify keywords and semantics in the voice: if the voice contains keywords such as "where to put it", "how to put it", or "cannot find", it is identified as a request for help; if it contains keywords such as "right", "is it", or "correct", it is identified as a confirmation; if it contains keywords such as "what is it", "what to eat", or "where", and points to a specific puzzle piece, it is identified as a knowledge question and answer; if it contains keywords such as "related to", "what is next to", or "what is together", and involves multiple puzzle pieces, it is identified as a cross-piece association.
[0035] The current user interaction intent is determined by comprehensively analyzing the interactive control commands and / or semantic analysis results obtained from the parsing. In other words, the user interaction intent is the core need for the user to initiate the interaction. It may be triggered by a single user input event or by a combination of multiple user input events. A comprehensive judgment based on both the interactive control commands and semantic analysis results is necessary to ensure the accuracy of intent recognition. Specifically, if only a pen gesture event is triggered, the user interaction intent is determined directly based on the interactive control command. For example, a single click corresponds to the intent to "play basic content," a double click corresponds to the intent to "explain in another way," and a long press corresponds to the intent to "seek help," etc. If only a voice input event is triggered, the user interaction intent is determined directly based on the semantic analysis results: the "seek help" semantic corresponds to the intent to "seek help," the "confirm" semantic corresponds to the intent to "confirm," the "knowledge question and answer" semantic corresponds to the intent to "knowledge query," and the "cross-block association" semantic corresponds to the intent to "cross-block association query." If both pen gesture events and voice input events are triggered simultaneously, the results of both are combined for a comprehensive judgment. For example, a long press combined with the voice prompt "Where to put it?" (the "seek help" semantic) determines the intent to be "location help." By analyzing various events, the current target puzzle piece can be identified and the user's interaction intent can be clarified, providing a more accurate input basis for determining the interaction state.
[0036] For example, if a user touches the "tree" puzzle piece and presses and holds the preset button on the reading pen while simultaneously inputting "Where do I put this?" (help semantics), the user's interaction intent is determined to be "help with the location of the tree" puzzle piece, i.e., the correct placement of the tree puzzle piece needs to be indicated; if the user clicks the reading pen to touch the "deer" puzzle piece (CMD1) without inputting any voice, the user's interaction intent is determined to be "play basic content".
[0037] In one embodiment, step S203 includes: Retrieve historical recognition records within the context window from the context cache data, and extract the historical recognition count and the most recent interaction record of the target puzzle piece from the historical recognition records based on the target puzzle piece information; Based on the historical recognition count of the target puzzle piece and the most recent interaction record, it is determined whether the current scenario is the first recognition scenario, the repeated recognition scenario, or the cross-piece switching scenario. Based on the user's interaction intent and the scenario confirmation result, the current corresponding interaction state is determined according to the preset state transition rules.
[0038] In this embodiment, all historical recognition records within the current context window are read from the context cache. The specific scope of the context window can be preset to a default value, such as the last 10 recognition events or the last 5 minutes, or it can be dynamically adjusted according to the user interaction frequency. These historical recognition records can be sorted in descending order by recognition timestamp. Then, based on the unique block ID of the target puzzle piece, all recognition records related to that block are filtered out and the recognition count is counted to obtain the historical recognition count of the target puzzle piece. At the same time, the record with the latest timestamp in the filtering results is extracted as the most recent interaction record, and information such as the interaction status, feedback content, and recognition time in that record is obtained. If there is no historical recognition record for the puzzle piece in the context window, the historical recognition count is determined to be 0, and the most recent interaction record is empty. By extracting the historical recognition count and the most recent interaction record of the target puzzle piece, historical basis is provided for the decision of the current interaction state, achieving contextual coherence.
[0039] Since the same puzzle piece is subject to different interaction strategies in the scenarios of first recognition, repeated recognition and cross-piece switching, when confirming the current interaction state to be entered, we first confirm whether the current scenario is the first recognition scenario, repeated recognition scenario or cross-piece switching scenario based on the historical recognition count of the target puzzle piece and the most recent interaction record, so as to output targeted puzzle interaction feedback. The specific criteria for classification are the number of historical recognitions, the ID of the most recently recognized puzzle piece, and the recognition time interval, ensuring the accuracy and rationality of the scenario classification. For example, the first recognition scenario refers to a scenario where the historical recognition count of the target puzzle piece is 0, or although there is a historical recognition record, the interval between the most recently recognized time and the current recognition time exceeds a preset threshold Tw1 (e.g., 5 minutes), and the historical context has expired. The repeated recognition scenario refers to a scenario where the unique block ID of the target puzzle piece is the same as the ID of the most recently recognized puzzle piece, and the time interval between the two recognitions is less than a preset threshold Tw2 (e.g., 30 seconds), or the number of recognitions of the block in the current context window is ≥2. The cross-block switching scenario refers to a scenario where the ID of the currently recognized target puzzle piece is different from the ID of the most recently recognized puzzle piece, and the time interval between the two recognitions is less than a preset threshold Tw3 (e.g., 1 minute), which is considered to be a block switching in the same continuous interaction link. Tw1, Tw2, and Tw3 can all be adjusted according to actual interaction needs, and this embodiment does not limit them.
[0040] Then, by combining the user's interaction intent and the scenario with the pre-maintained state transition rules, the current corresponding interaction state is determined. These pre-set state transition rules clarify the target interaction state corresponding to different scenarios and different user interaction intents, ensuring the orderliness and rationality of state switching. For example, in the first recognition scenario, if there is no clear interaction intent, it corresponds to the first explanation state; if there is a knowledge query intent, it corresponds to the knowledge question and answer state; if there is a request for help, it corresponds to the location request for help state, etc. In the repeated recognition scenario, if there is no clear interaction intent, it corresponds to the repeated recognition review state; if there is a request for help, it corresponds to the location request for help state, etc. In the cross-block switching scenario, if there is a cross-block association intent, it corresponds to the cross-block association state; if there is no clear interaction intent, it corresponds to the first explanation state (basic explanation of the new block), etc. The specific state transition rules can be flexibly adjusted according to actual needs, and this embodiment does not limit them.
[0041] This embodiment divides different interaction scenarios and then determines the interaction state based on user intent, avoiding the subjectivity and randomness of state determination, ensuring that the interaction state is highly adapted to user needs and interaction scenarios, and further improving the adaptive capability of interaction state switching.
[0042] In one embodiment, determining the current interaction state according to a preset state transition rule based on the user interaction intent and the scenario confirmation result includes: The basic interaction state is determined based on the scenario confirmation result, wherein the first recognition scenario corresponds to the first narration state, the repeated recognition scenario corresponds to the repeated recognition review state, and the cross-block switching scenario corresponds to the cross-block association state. According to the user's interaction intent, the basic interaction state is transitioned according to the preset state transition rules to confirm the candidate interaction state. The candidate interaction states are prioritized, and the candidate interaction state with the highest priority is determined as the current interaction state.
[0043] In this embodiment, the basic interaction state is the default interaction state determined based on the scene confirmation result. Each scene corresponds to a core basic interaction state. The first recognition scene corresponds to the first explanation state, which introduces the basic knowledge points of the target puzzle piece to the user and helps the user to understand the puzzle piece initially. The repeated recognition scene corresponds to the repeated recognition review state, which reviews and supplements the knowledge points that have been explained, avoids repeating basic content, and improves interaction efficiency. The cross-piece switching scene corresponds to the cross-piece association state, which establishes the association between the current puzzle piece and the most recently recognized puzzle piece, helping the user understand the overall logic and relationship of the puzzle. By determining the basic interaction state, a benchmark is provided for subsequent state transitions, ensuring the continuity and rationality of the interaction state and avoiding confusion in the interaction state.
[0044] Then, based on the user's interaction intent and according to preset state transition rules, the basic interaction state is transitioned to identify candidate interaction states. These candidate interaction states are the possible states after the basic interaction state transitions according to the user's interaction intent; there may be one or more such states. The transition logic is determined by the preset state transition rules, thus adapting to the user's specific interaction needs. Furthermore, if only one candidate interaction state exists, it is directly determined as the current interaction state; if multiple candidate states exist, they are prioritized. This ensures that when multiple candidate states exist, the interaction logic that best meets the user's current needs and is most urgent is executed first, avoiding interaction chaos caused by multiple state conflicts. The priority ranking rules can be preset based on the urgency, educational value, and user needs of the interaction to ensure the rationality of the ranking.
[0045] In practical implementation, state machine control can be used to switch between multiple interactive states. For example, the state machine maintains at least the following states: S0 Idle state, S1 Initial presentation state, S2 Repeated recognition and review state, S3 Problem understanding state, S4 Knowledge Q&A state, S5 Location assistance state, S6 Error correction confirmation state, S7 Encouraging feedback state, S8 Cross-block association state, S9 Timeout recovery state, and the state transition follows these rules: After receiving a block identification event from S0, if it is the first identification, proceed to S1; if it is a repeated identification, proceed to S2. After receiving a voice question event from S1 / S2, proceed to S3; In S3, if a help-seeking semantic is identified, proceed to S5; if a confirmation semantic is identified, proceed to S6; if a knowledge question is identified, proceed to S4; if a cross-block association semantic is identified, proceed to S8. In S5, if the user has not yet completed the stitching and requests help again, S5 is maintained and the prompt level is raised; if the user completes the stitching, the process proceeds to S7. In S6, if the placement is confirmed to be incorrect, proceed to S5 or S7; if it is confirmed to be correct, return to S1 or end. If a timeout occurs in any state, proceed to S9; in S9, if the user presses the same block again, the state before the timeout will be restored; if a new block is pressed, proceed to S1 or S2 for the corresponding new block. After outputting in S4, S5, S6, and S8, if the system determines that encouragement is needed, it will enter S7 and then return to the waiting state.
[0046] Preferably, the state priority is sorted according to "location help > error correction confirmation > problem understanding / knowledge Q&A > cross-block association > repeated recognition review > first explanation > encouragement feedback". If multiple conditions are hit in the same round of input, the state switch with the higher priority is executed.
[0047] In one embodiment, after step S203, the method further includes: Confirm the preset timeout threshold based on the current interaction status; Detect whether a timeout event has occurred in the current interaction state based on the timeout threshold; If a timeout event occurs, the corresponding recovery prompt will be executed based on the current interaction state.
[0048] In this embodiment, in addition to puzzle piece recognition events and user input events, timeout events are also identified during the interaction process. The timeout threshold refers to the maximum time to wait for the user's next interaction in the current interaction state. Different interaction states have different timeout thresholds, specifically the reasonable time required for the user to complete the interaction in that state, avoiding misjudgments caused by a uniform timeout threshold. Specifically, a corresponding timeout threshold is pre-set for each interaction state: Tn1 = 10 seconds for the initial explanation state (time required for the user to listen to the basic explanation); Tn2 = 15 seconds for the knowledge question and answer state (time required for the user to think and answer the question); Tn3 = 20 seconds for the location assistance state (time required for the user to attempt to piece the puzzle based on location prompts); Tn4 = 8 seconds for the repeated recognition review state; Tn5 = 12 seconds for the cross-piece association state; Tn6 = 10 seconds for the error correction confirmation state, and so on. After determining the current interaction state, the timeout threshold corresponding to that state is retrieved as the basis for subsequent timeout detection. Furthermore, the timeout threshold can be dynamically adjusted according to the user group. For example, the timeout threshold is different for children of different ages, with younger children having a longer timeout threshold, in order to improve adaptability.
[0049] Based on the acquired timeout threshold, the system detects whether a timeout event has occurred in the current interaction state. Specifically, after outputting the feedback content of the current interaction state, a timer is started to begin timing, while simultaneously listening for user input events. If a user input event is received during the timing process, the timer is immediately stopped, reset, and the next round of the interaction process begins. If the timing time reaches the preset timeout threshold and no user input event is received, a timeout event is determined to have occurred, triggering the timeout handling process. By detecting timeout events, user interaction stagnation can be identified in a timely manner, preventing the system from being in a waiting state for an extended period and improving the smoothness of the interaction.
[0050] Upon detecting a timeout event, the system further executes a corresponding recovery prompt based on the current interaction state. This recovery prompt means that after a timeout event occurs, targeted prompt content is output based on the current interaction state to guide the user to continue interacting. At the same time, the current context information is saved to avoid context loss, ensure the continuity of interaction, and improve the user experience. Different interaction states correspond to different recovery prompts. For example, if the current state is the initial narration, a light follow-up question or encouraging prompt will be output when a timeout occurs; if the current state is the knowledge Q&A state, a simplified restatement or clarification prompt will be output when a timeout occurs; if the current state is the location help state, the prompt level will be automatically raised by one level when a timeout occurs, or the user will be asked to press the relevant puzzle piece again, such as "Just a reminder, the deer is next to the big tree" or "Please press the deer puzzle piece again to get more prompts"; if two consecutive timeouts occur, the context of the current piece will be downgraded to "context to be restored". When the user presses the same piece again, it will be restored to the previous prompt level. When the user presses a new piece, the short-term state of the piece will be cleared, and only the topic-level context will be retained.
[0051] This embodiment uses a timeout detection and recovery prompt mechanism to set differentiated timeout thresholds based on different interaction states. This allows for timely detection of user interaction stagnation, preventing users from being in a waiting state for extended periods. At the same time, targeted recovery prompts guide users to continue interacting, ensuring the continuity of the interaction.
[0052] In one embodiment, step S204 includes: The corresponding feedback template is invoked based on the current interaction state, and the set of output content of the target puzzle piece is obtained from the context cache data; Based on the set of output content, confirm the uncovered slots in the feedback module; Based on the target puzzle piece information and the user's interaction intent, generate the filling content for the uncovered slots and output the corresponding puzzle interaction feedback; After updating the current interaction state and the output puzzle interaction feedback to the context cache, return to the listening state and wait for the next round of event triggering.
[0053] In this embodiment, when outputting puzzle interaction feedback in the current interaction state, the corresponding feedback template is called from a preset feedback template library according to the current interaction state. This feedback template is a preset feedback content framework corresponding to the interaction state, including fixed phrases and fillable slots, etc. Different interaction states correspond to different feedback templates to ensure the standardization and relevance of the feedback content, while improving the efficiency of feedback content generation. At the same time, the set of output content for the target puzzle piece is extracted from the context cache data, that is, the set of feedback content that has been output for the current target puzzle piece in the context window, including the set of output content slots, the set of covered knowledge points, the set of used prompt levels, and the set of questions that the user has answered correctly, etc.
[0054] The system matches the content in the already output content set with each slot in the feedback template, determining whether the content corresponding to each slot has already been output. If the content corresponding to a slot already exists in the already output content set, the slot is determined to be an covered slot; otherwise, it is an uncovered slot. Based on the target puzzle piece information and the user's interaction intent, the system generates fill content for the uncovered slots and finally outputs the corresponding puzzle interaction feedback. By confirming the uncovered slots, the system ensures that the generated feedback content only outputs information that has not been explained before, avoiding repeated output and improving interaction efficiency.
[0055] Before generating a new response, the system first checks which content slots corresponding to the current target puzzle piece have already been output and which have not yet been covered. Preferably, if the basic explanation slot has already been output, the next time the system will not repeat the complete explanation, but will only output the supplementary knowledge points that have not been covered; if the knowledge point has been output and the user has answered correctly, the knowledge point will be omitted in subsequent questions and the system will move on to a higher-level question; if the user did not understand or answered incorrectly in the previous question, a simplified version of the knowledge point will be restated first. Through content slot deduplication, knowledge point coverage detection, and answer status recording, intelligent control is achieved to avoid repeating explanations, omit redundant information, and supplement content not covered in the previous text.
[0056] In one embodiment, when the user input event includes a voice input event, the method further includes: A block-level context window is established based on the target puzzle piece information. Voice input events are received within the duration of the upper and lower windows of the block level, and a binding relationship is established between the voice input events and the target puzzle piece. If an object name is parsed in the voice input event, the matching degree between the object name and the target puzzle piece and adjacent puzzle pieces is calculated. The degree of association between the voice input event and the target puzzle piece is determined based on the matching degree. The binding relationship is dynamically adjusted within the block-level context window based on the degree of association, and corresponding adjustment prompts are output.
[0057] In this embodiment, when user input events include voice input events, voice questions are also bound to physical puzzle pieces to avoid generalized answers that are out of context. Specifically, after receiving a puzzle piece recognition event and obtaining the target puzzle piece information, a block-level context window is established and a window duration Tq is set. Voice input events received within this context window are bound to the target puzzle piece to ensure that subsequent semantic parsing and feedback can accurately associate with the puzzle piece. That is, after the user presses a puzzle piece, a block-level context window is established. Within the preset duration Tq, if the user asks a voice question and does not switch the block ID again, the voice question is bound to the current block by default. If no voice input event is received within the block-level context window, the window is automatically closed without establishing any binding relationship. By establishing a block-level context window to limit the association range between voice input events and target puzzle pieces, the association of voice input events unrelated to the current puzzle piece is avoided, achieving a strong association between voice interaction and physical puzzle pieces. This ensures that the system can clearly identify the puzzle piece corresponding to the voice question and avoids giving irrelevant answers.
[0058] At this point, if pronouns such as "it," "this," or "here" are parsed from the voice input event, they are preferentially interpreted as referring to the current puzzle piece. If a specific object name is parsed from the voice input event, the matching degree between that object name and the target puzzle piece, as well as adjacent puzzle pieces, is calculated. Based on the matching degree, the degree of association between the voice input event and the target puzzle piece is determined. For example, the degree of association can be divided into three categories: "high association," "medium association," and "low association," corresponding to different matching degree ranges. The higher the matching degree, the higher the degree of association. Based on the degree of association, the binding relationship within the context window is adjusted, thereby adjusting the binding object between the voice input event and the puzzle piece to ensure the accuracy of the binding relationship. At the same time, corresponding adjustment prompts are output to let the user clearly understand the puzzle piece corresponding to the voice question, improving the transparency of the interaction and the user experience.
[0059] Specifically, if the question contains a clear object name, the matching degree between that object name and the semantic tag of the current block, the tags of adjacent blocks, and the current topic tag is calculated. When the matching degree is higher than the threshold R1 (e.g., 0.8), it is confirmed that the voice input event is highly related to the target puzzle piece, and the question is processed according to the current target puzzle piece. When the matching degree falls in the middle range (e.g., 0.3-0.8), it is confirmed as moderately related, and the clarification process is initiated with a prompt message, such as "Are you asking about this deer piece, or the whole forest puzzle?". When the matching degree is lower than the threshold R2 (e.g., 0.3), it is confirmed as lowly related. In this case, the current puzzle piece is not used directly to answer; the original binding relationship can be removed, and the user can be prompted to press the relevant puzzle piece again, or it can be explained that the current question is not sufficiently related to the puzzle scene. If a low degree of correlation is confirmed while there is an adjacent puzzle piece with a higher matching degree, the binding relationship is adjusted to that adjacent puzzle piece, and the prompt message "Are you asking about XX next to you? Here's the answer," thus avoiding erroneous responses to questions unrelated to the current puzzle piece.
[0060] In the above embodiments, this invention discloses an adaptive interaction method based on jigsaw puzzle piece recognition. The method receives jigsaw puzzle piece recognition events and user input events, including pen action events and / or voice input events. It parses the jigsaw puzzle piece recognition events to obtain target jigsaw puzzle piece information and parses the user input events to obtain user interaction intent. It retrieves historical recognition records from a context window and determines the current interaction state based on the target jigsaw puzzle piece information, user interaction intent, and historical recognition records. Based on the current interaction state, it outputs corresponding jigsaw puzzle interaction feedback and updates the context cache data. This invention overcomes the problem of traditional jigsaw puzzle teaching aids only playing fixed audio and lacking scene-adaptive interaction by recognizing jigsaw puzzle piece recognition events and user input events, combined with context history records to achieve intelligent switching of interaction states. This effectively improves the continuity and intelligence level of jigsaw puzzle reading interaction.
[0061] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0062] Further reference Figure 3 As a response to the above Figure 2 The present invention provides an embodiment of an adaptive interactive device based on puzzle piece recognition, which is similar to the method shown. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0063] like Figure 3As shown, the adaptive interactive device 30 based on puzzle piece recognition described in this embodiment includes: Event receiving module 301 is used to receive puzzle piece recognition events to be processed and user input events, wherein the user input events include pen tip action events and / or voice input events; The event parsing module 302 is used to parse the puzzle piece recognition event to obtain target puzzle piece information, and to parse the user input event to obtain user interaction intent; The interaction state control module 303 is used to obtain historical recognition records in the context window and determine the current interaction state based on the target puzzle piece information, user interaction intent and historical recognition records. The interactive feedback module 304 is used to output corresponding puzzle interactive feedback based on the current interactive state and update the context cache data.
[0064] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the adaptive interactive execution process based on jigsaw puzzle recognition. For specific implementation methods of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0065] In one embodiment, the event parsing module 302 includes: The puzzle piece identification extraction unit is used to parse the puzzle piece identification event and extract the unique block ID of the target puzzle piece; The action parsing unit is used to parse the pen tip action events, identify the pen tip action type, and encode it into corresponding interactive control commands; The semantic parsing unit is used to perform semantic parsing on voice input events and identify help-seeking semantics, confirmation semantics, knowledge-based question-and-answer semantics, and / or cross-block association semantics in the current input voice. The intent determination unit is used to determine the user's interaction intent based on the interaction control instructions and / or semantic parsing results.
[0066] In one embodiment, the interactive state control module 303 includes: The historical information acquisition unit is used to acquire historical recognition records within the context window from the context cache data, and extract the historical recognition count and the most recent interaction record of the target puzzle piece from the historical recognition records based on the target puzzle piece information; The scene recognition unit is used to confirm whether the current scene is the first recognition scene, the repeated recognition scene, or the cross-block switching scene based on the historical recognition count of the target puzzle piece and the most recent interaction record. The state transition control unit is used to determine the current corresponding interaction state according to the user's interaction intent and the scenario confirmation result, based on a preset state transition rule.
[0067] In one embodiment, the state transition control unit includes: The basic confirmation unit is used to determine the basic interaction state based on the scenario confirmation result, wherein the first recognition scenario corresponds to the first narration state, the repeated recognition scenario corresponds to the repeated recognition review state, and the cross-block switching scenario corresponds to the cross-block association state. The state transition unit is used to perform state transition on the basic interaction state according to the user interaction intent and a preset state transition rule, and to confirm the candidate interaction state. The sorting and confirmation unit is used to sort the candidate interaction states by priority and determine the candidate interaction state with the highest priority as the current interaction state.
[0068] In one embodiment, the device 30 further includes: The threshold confirmation module is used to confirm the preset timeout threshold based on the current interaction state. The timeout detection module is used to detect whether a timeout event has occurred in the current interaction state based on the timeout threshold. The timeout response module is used to perform corresponding recovery prompts based on the current interaction state if a timeout event occurs.
[0069] In one embodiment, the interactive feedback module 304 includes: The data retrieval unit is used to invoke the corresponding feedback template based on the current interaction state and obtain the set of output content of the target puzzle piece from the context cache data; The slot confirmation unit is used to confirm the uncovered slots in the feedback module based on the set of output content. The feedback output unit is used to generate the filling content of the uncovered slots based on the target puzzle piece information and the user's interaction intent, and output the corresponding puzzle interaction feedback; The cache update unit is used to update the current interaction state and the output puzzle interaction feedback to the context cache and then return to the listening state to wait for the next round of event triggering.
[0070] In one embodiment, the device 30 further includes: The binding module is used to establish a block-level context window based on the target puzzle piece information, receive voice input events within the duration of the upper and lower windows of the block level, and establish a binding relationship between the voice input events and the target puzzle piece. The matching module is used to calculate the matching degree between the object name and the target puzzle piece and adjacent puzzle pieces if the object name is parsed in the voice input event; The association confirmation module is used to confirm the degree of association between the voice input event and the target puzzle piece based on the matching degree. The binding adjustment module is used to dynamically adjust the binding relationship within the block-level context window according to the degree of association, and output corresponding adjustment prompts.
[0071] In the above embodiments, this invention discloses an adaptive interactive device based on jigsaw puzzle piece recognition. The device receives jigsaw puzzle piece recognition events and user input events, including pen action events and / or voice input events. It parses the jigsaw puzzle piece recognition events to obtain target jigsaw puzzle piece information and parses the user input events to obtain the user's interaction intent. It retrieves historical recognition records from a context window and determines the current interaction state based on the target jigsaw puzzle piece information, the user's interaction intent, and the historical recognition records. Based on the current interaction state, it outputs corresponding jigsaw puzzle interaction feedback and updates the context cache data. This invention overcomes the problem of traditional jigsaw puzzle teaching aids only playing fixed audio and lacking scene-adaptive interaction by recognizing jigsaw puzzle piece recognition events and user input events, combined with context history records to achieve intelligent switching of interaction states. This effectively improves the continuity and intelligence level of jigsaw puzzle reading interaction.
[0072] Specific limitations regarding the adaptive interactive device based on jigsaw puzzle recognition can be found in the limitations of the adaptive interactive method based on jigsaw puzzle recognition described above, and will not be repeated here. Each module in the aforementioned adaptive interactive device based on jigsaw puzzle recognition can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0073] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0074] Another embodiment of the present invention provides a computer device, such as... Figure 4 As shown, computer device 40 includes: One or more processors 401 and memory 402, Figure 4 The following section uses a processor 401 as an example. The processor 401 and the memory 402 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.
[0075] The processor 401 is used to perform various control logics of the computer device 40. It can be any conventional processor, microprocessor, state machine, general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0076] The memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the adaptive interaction method based on jigsaw puzzle recognition in the embodiments of the present invention. The processor 401 executes various functional applications and data processing of the computer device 40 by running the non-volatile software programs, instructions, and units stored in the memory 402, thereby implementing the adaptive interaction method based on jigsaw puzzle recognition in the above method embodiments.
[0077] Another embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, perform the steps of the adaptive interaction method based on puzzle piece recognition in any of the above method embodiments.
[0078] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0079] Based on the above description of the embodiments, those skilled in the art will understand that the methods described in the embodiments can be implemented using software plus necessary general-purpose hardware platforms. Of course, they can also be implemented using hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0080] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0081] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0082] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An adaptive interaction method based on puzzle piece recognition, characterized in that, include: Receive puzzle piece recognition events and user input events to be processed, wherein the user input events include pen tip action events and / or voice input events; The puzzle piece recognition event is parsed to obtain the target puzzle piece information, and the user input event is parsed to obtain the user's interaction intent; Obtain historical recognition records from the context window, and determine the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records; Output the corresponding puzzle interaction feedback based on the current interaction state, and update the context cache data; When the user input event includes a voice input event, the method further includes: A block-level context window is established based on the target puzzle piece information. Voice input events are received within the duration of the block-level context window, and a binding relationship is established between the voice input events and the target puzzle piece. If an object name is parsed in the voice input event, the matching degree between the object name and the target puzzle piece and adjacent puzzle pieces is calculated. The degree of association between the voice input event and the target puzzle piece is determined based on the matching degree. The binding relationship is dynamically adjusted within the block-level context window based on the degree of association, and corresponding adjustment prompts are output.
2. The adaptive interaction method based on puzzle piece recognition according to claim 1, characterized in that, The process of parsing the puzzle piece recognition event to obtain target puzzle piece information and parsing the user input event to obtain user interaction intent includes: The puzzle piece recognition event is parsed to extract the unique block ID of the target puzzle piece; The pen tip action events are parsed to identify the pen tip action type and encode it into corresponding interactive control commands; Perform semantic parsing on voice input events to identify help-seeking semantics, confirmation semantics, knowledge-based question-and-answer semantics, or cross-block association semantics in the current input voice; The user's interaction intent is determined based on the interaction control instructions and / or semantic parsing results.
3. The adaptive interaction method based on puzzle piece recognition according to claim 1, characterized in that, The step of obtaining historical recognition records in the context window and determining the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records includes: Retrieve historical recognition records within the context window from the context cache data, and extract the historical recognition count and the most recent interaction record of the target puzzle piece from the historical recognition records based on the target puzzle piece information; Based on the historical recognition count of the target puzzle piece and the most recent interaction record, it is determined whether the current scenario is the first recognition scenario, the repeated recognition scenario, or the cross-piece switching scenario. Based on the user's interaction intent and the scenario confirmation result, the current corresponding interaction state is determined according to the preset state transition rules.
4. The adaptive interaction method based on puzzle piece recognition according to claim 3, characterized in that, The step of determining the current interaction state according to the user's interaction intent and the scenario confirmation result, based on a preset state transition rule, includes: The basic interaction state is determined based on the scenario confirmation result, wherein the first recognition scenario corresponds to the first narration state, the repeated recognition scenario corresponds to the repeated recognition review state, and the cross-block switching scenario corresponds to the cross-block association state. According to the user's interaction intent, the basic interaction state is transitioned according to the preset state transition rules to confirm the candidate interaction state. The candidate interaction states are prioritized, and the candidate interaction state with the highest priority is determined as the current interaction state.
5. The adaptive interaction method based on puzzle piece recognition according to claim 1, characterized in that, After obtaining the historical recognition records in the context window and determining the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records, the method further includes: Confirm the preset timeout threshold based on the current interaction status; Detect whether a timeout event has occurred in the current interaction state based on the timeout threshold; If a timeout event occurs, the corresponding recovery prompt will be executed based on the current interaction state.
6. The adaptive interaction method based on puzzle piece recognition according to claim 1, characterized in that, The process of outputting corresponding puzzle interaction feedback based on the current interaction state and updating context cache data includes: The corresponding feedback template is invoked based on the current interaction state, and the set of output content of the target puzzle piece is obtained from the context cache data; Based on the set of output content, confirm the uncovered slots in the feedback template; Based on the target puzzle piece information and the user's interaction intent, generate the filling content for the uncovered slots and output the corresponding puzzle interaction feedback; After updating the current interaction state and the output puzzle interaction feedback to the context cache data, return to the listening state and wait for the next round of event triggering.
7. An adaptive interactive device based on jigsaw puzzle piece recognition, characterized in that, include: The event receiving module is used to receive puzzle piece recognition events and user input events to be processed, wherein the user input events include pen tip action events and / or voice input events; The event parsing module is used to parse the puzzle piece recognition event to obtain target puzzle piece information, and to parse the user input event to obtain user interaction intent; The interaction state control module is used to obtain historical recognition records in the context window and determine the current interaction state based on the target puzzle piece information, user interaction intent, and historical recognition records. The interactive feedback module is used to output corresponding puzzle interaction feedback based on the current interaction state and update the context cache data; The binding module is used to establish a block-level context window based on the target puzzle piece information when the user input event includes a voice input event, receive the voice input event within the duration of the block-level context window, and establish a binding relationship between the voice input event and the target puzzle piece. The matching module is used to calculate the matching degree between the object name and the target puzzle piece and adjacent puzzle pieces if the object name is parsed in the voice input event; The association confirmation module is used to confirm the degree of association between the voice input event and the target puzzle piece based on the matching degree. The binding adjustment module is used to dynamically adjust the binding relationship within the block-level context window according to the degree of association, and output corresponding adjustment prompts.
8. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the adaptive interaction method based on puzzle piece recognition as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the adaptive interaction method based on puzzle piece recognition as described in any one of claims 1-6.
Citation Information
Patent Citations
Learning content output method and learning equipment
CN111081092A