Portable cognitive device based on artificial intelligence

By combining optical positioning marking technology of the image acquisition module and the lighting projection module, the portability and interactivity of existing devices are solved, and efficient, safe and intelligent reading assistance for screen-free portable cognitive devices are realized.

CN120526432APending Publication Date: 2025-08-22SHENZHEN PERCHERRY TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510636592.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-17
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing children's picture book reading auxiliary equipment requires pre-construction of a huge image-voice database. The equipment is large and the screen is easy to distract, which affects the convenience of use and the accuracy of recognition.

Method used

The lightweight image acquisition module is combined with the lighting projection module, and the effective acquisition area is indicated through optical positioning marks, combined with intelligent voice wake-up and intention recognition mechanisms to achieve multimodal interaction, and real-time processing and security audits are carried out through AI services.

Benefits of technology

It breaks through the limitations of traditional devices relying on pre-stored data, reduces cost and volume, improves acquisition accuracy and user attention, and realizes automated guidance and secure multimodal interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526432A_ABST
    Figure CN120526432A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to portable cognitive equipment based on artificial intelligence. According to the technical scheme, the light-weight image acquisition module is combined with the lamplight projection module, and the effective acquisition area is indicated by the optical positioning mark, so that the portable cognitive equipment without a screen is formed; during specific implementation, a user aligns equipment to a picture book, the lamplight projection module projects an optical positioning mark to indicate a view finding area, and after an acquisition signal is triggered, a target image can be acquired and processed in real time through AI service; the limitation that traditional equipment depends on pre-stored data is broken through, the cost and the size are remarkably reduced by removing a screen display module, the attention of a user is improved, and meanwhile the technical problem of accurate view finding is solved through the projection light frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a portable cognitive device based on artificial intelligence. Background Art

[0002] With the development of artificial intelligence technology, the application of smart devices in assisting reading and cognition is becoming increasingly widespread. In the field of children's education, in particular, using smart devices to assist reading has become an important educational tool, helping children better understand and remember reading content, thereby improving learning outcomes.

[0003] Traditional children's picture book reading systems typically use cartoon dolls, learning tablets, or children's cameras equipped with cameras. To use the system, the picture book is placed under the camera, and the system uses image matching to retrieve the corresponding audio content from a pre-stored database and play it back, enabling the picture book's reading function.

[0004] However, existing equipment requires the collection of a large amount of picture book data in advance, the number of picture books that can be read is limited, and the device screen can easily distract children, making it difficult to concentrate on reading. This situation needs further improvement. Summary of the Invention

[0005] To address the problems that existing devices require the pre-collection of large amounts of picture book data and that the device screen easily distracts children, hindering their focused reading, this application provides a portable cognitive device based on artificial intelligence, including: case; An image acquisition module, disposed in the housing, for acquiring a target image; a light projection module, disposed in the housing, for projecting an optical positioning mark indicating a viewing area into the field of view of the image acquisition module; a processor, disposed in the housing; A sound playing module is disposed in the housing; a communication module, disposed in the housing; Among them, the optical positioning mark is used to indicate the effective acquisition area of ​​the image acquisition module, the processor is used to control the image acquisition module to acquire the target image and process it, and send the processing results to the server through the communication module for subsequent processing.

[0006] By adopting the above-mentioned technical solution, the present application solves the technical problems of existing children's picture book reading assistance devices; traditional picture book reading assistance devices, such as story robots or learning tablets, require the pre-construction of a huge image-voice database and the large size of the devices, which seriously affects the convenience of use; the present application combines a lightweight image acquisition module with a light projection module, and adopts a technical solution of optical positioning marks to indicate the effective acquisition area to form a screenless portable cognitive device; in specific implementation, the user points the device at the picture book, and the light projection module projects an optical positioning mark to indicate the framing area. After triggering the acquisition signal, the target image can be acquired and processed in real time through local or cloud AI services; it not only breaks through the limitation of traditional devices relying on pre-stored data, but also significantly reduces the cost and volume by removing the screen display module, improves the user's attention, and at the same time solves the technical problem of accurate framing by using the projected light frame.

[0007] Optionally, the processor is further configured to: Detecting the position of the optical positioning mark in the target image to determine the effective acquisition area; When the position of the optical positioning mark cannot be detected, determining whether the distance between the device and the target is too close, and providing a prompt through the sound playing module; Detecting a text area within the effective acquisition area; Based on the distribution position of the text area, the page type is identified, and the page type includes a cover page, a text page or other pages.

[0008] By adopting the above-mentioned technical solution, the present application solves the technical problem that existing portable cognitive devices are difficult to accurately locate the collection area during actual use; since traditional devices lack an effective spatial positioning mechanism, users often need to repeatedly adjust the device position and angle when collecting images, which seriously affects the user experience and recognition accuracy; the present application detects the position of the optical positioning mark in the target image and combines it with the distance judgment mechanism. The device first detects the position of the optical positioning mark in the target image to determine the effective collection area. When the cursor center position cannot be detected, it automatically judges the distance between the device and the target and gives a voice prompt. Then, it detects the text distribution in the determined effective area and intelligently identifies the page type; it not only solves the problem of inaccurate positioning of traditional devices, but also improves the collection accuracy through intelligent distance detection and page type recognition, and at the same time realizes automated guidance of the collection process.

[0009] Optionally, also include: A sound collection module is provided in the housing; The processor is further configured to: caching the target image captured by the image acquisition module; Detecting whether the voice input of the sound collection module contains a preset wake-up word; When the wake-up word is detected, converting the voice input into text content and analyzing the text content to determine user intent; selectively combining the cached target image with the text content according to the user intention; When the input of the sound collection module exceeds a preset time, the detection is stopped and a prompt is given through the sound playing module.

[0010] By adopting the above-mentioned technical solution, the present application solves the technical problem that existing portable cognitive devices are single and lack intelligence in human-computer interaction; since traditional devices usually only support simple button operations or touch controls, it is difficult for users to achieve complex interaction needs, especially in scenarios where image and voice input need to be processed simultaneously, the performance is poor; the present application proposes a multimodal interaction solution by integrating a sound acquisition module and combining it with an intelligent voice wake-up and intention recognition mechanism; in specific implementation, the device will automatically cache the collected target image, and at the same time monitor whether the voice input contains a preset wake-up word. When the wake-up word is detected, the voice will be converted into text and the user's intention will be analyzed, and then the image and text content will be intelligently combined according to the intention. It can also actively stop detection and prompt the user when the voice input times out; the flexibility of use is improved through intelligent voice control and content combination mechanism, and the collaborative processing of images and voice is realized.

[0011] Optionally, the processor is further configured to: Preprocessing the target image, wherein the preprocessing includes image scaling and color space conversion; Recognizing QR code content in the target image; If a QR code is identified, it is sent to an artificial intelligence service to determine whether the URL in the QR code is safe and compliant, filter illegal URLs, and analyze and filter the content pointed to by the QR code. extracting playable content from the processing results of the artificial intelligence service; When illegal content is detected, a warning is issued through the sound playing module.

[0012] By adopting the above-mentioned technical solution, the present application solves the security risks that exist in existing portable cognitive devices when processing QR code content. Since traditional devices usually jump directly to or play related content after identifying the QR code, they lack an effective security review mechanism, which makes it easy for users, especially children, to be exposed to bad information or suffer from online fraud. The present application establishes a complete image preprocessing and QR code content security detection mechanism. The device first performs preprocessing such as scaling and color space conversion on the target image, and then identifies the QR code content and sends it to the artificial intelligence service for a comprehensive security compliance check, including URL security verification, illegal content filtering, etc. Finally, only content that can be safely played is extracted, and timely warnings are issued when potential risks are detected. The intelligent analysis of the AI ​​service provides a more reliable protection mechanism, while realizing real-time security monitoring and early warning.

[0013] Optionally, the processor is further configured to: Establish a user reading behavior model and record reading habit data including device movement frequency, average reading time, and page turning interval; Predicting the type and reading difficulty of the next page based on the reading habit data; Based on the prediction results, the display parameters of the optical positioning mark and the voice playback speed are adaptively adjusted.

[0014] By adopting the above-mentioned technical solution, this application establishes a user reading behavior model, uses data such as device movement frequency, average reading time and page turning interval to predict page type and reading difficulty, and then dynamically adjusts the optical positioning mark display parameters and voice playback speed, significantly improving reading efficiency and experience.

[0015] Optionally, the processor is further configured to: Detect ambient light intensity and ambient noise level; evaluating the suitability of the current reading environment based on the reading habit data; When the environmental suitability is lower than a preset threshold, the brightness of the optical positioning mark and the volume of the sound playing module are adjusted to help the user maintain attention.

[0016] By adopting the above-mentioned technical solution, this application monitors the ambient light intensity and noise level in real time, combines user reading habit data to build an environmental suitability assessment model, and dynamically adjusts the brightness of the optical positioning mark and voice playback parameters, breaking through the limitations of the passive response of traditional devices, achieving active environmental adaptation, and significantly improving reading concentration and natural interaction in complex scenarios.

[0017] Optionally, the processor is further configured to: identifying the user's reading proficiency based on the reading habit data; The optimal acquisition distance range between the device and the target, the tracking sensitivity of the optical positioning marker, and the interactive frequency of the voice playback are dynamically adjusted according to the reading proficiency.

[0018] By adopting the above technical solution, this application deeply mines the implicit features in reading habit data, builds a user reading proficiency evaluation system, dynamically optimizes the device acquisition distance, light frame tracking sensitivity and voice interaction frequency, and solves the interaction efficiency bottleneck problem of reading assistive devices caused by differences in user abilities.

[0019] Optionally, the processor is further configured to: Statistically analyze the length of time users stay on different types of pages and the number of times they read them repeatedly to obtain statistical analysis results; establishing a personalized reading difficulty assessment model based on the statistical analysis results; Based on the reading difficulty assessment model, the dwell time of the optical positioning mark is adjusted and the speech speed and number of repetitions of the voice playback are optimized.

[0020] By adopting the above-mentioned technical solution, this application constructs a dynamic reading difficulty assessment model through multimodal behavioral data analysis, and uses the user's stay time and repeated reading times on different types of pages as core features to achieve personalized adjustment of the optical positioning mark stay time and voice parameters, significantly improving reading comprehension efficiency and cognitive load adaptability in complex scenarios.

[0021] Optionally, the processor is further configured to: Detecting the layout features of the text area, dividing the optical positioning mark into multiple sub-areas, and adaptively adjusting the size of each sub-area according to the cognitive complexity of the text content; For the content in each sub-area, identify the unfamiliarity and grammatical structure of the text and determine the corresponding cognitive difficulty; The dwell time and reading speed of each sub-area are adjusted according to the cognitive difficulty.

[0022] By adopting the above-mentioned technical solution, this application deeply analyzes the typesetting features and semantic content of the text area, constructs a dynamic sub-area cognitive difficulty model, realizes hierarchical adaptive adjustment of optical positioning marks, and significantly improves the efficiency of complex text processing and the utilization of user cognitive resources.

[0023] In summary, this application includes at least one of the following beneficial technical effects: This application solves the technical problems of existing children's picture book reading assistance devices. Traditional picture book reading assistance devices, such as story robots or learning tablets, require the pre-construction of a huge image-speech database and are large in size, which seriously affects the convenience of use. This application combines a lightweight image acquisition module with a light projection module and adopts a technical solution of optical positioning marks to indicate the effective acquisition area to form a screenless portable cognitive device. During specific implementation, the user points the device at the picture book, and the light projection module projects an optical positioning mark to indicate the framing area. After triggering the acquisition signal, the target image can be acquired and processed in real time through AI services. This not only breaks through the limitation of traditional devices relying on pre-stored data, but also significantly reduces cost and size by removing the screen display module, improves user attention, and solves the technical problem of accurate framing by using the projected light frame. Since traditional devices lack an effective spatial positioning mechanism, users often need to repeatedly adjust the device position and angle when performing image acquisition, which seriously affects the user experience and recognition accuracy. The present application detects the position of the optical positioning mark in the target image and combines it with a distance judgment mechanism. The device first detects the position of the optical positioning mark in the target image to determine the effective acquisition area. When the cursor center position cannot be detected, it automatically determines the distance between the device and the target and gives a voice prompt. Then, it detects the text distribution in the determined effective area and intelligently identifies the page type. This not only solves the problem of inaccurate positioning of traditional devices, but also improves the acquisition accuracy through intelligent distance detection and page type recognition, while realizing automated guidance of the acquisition process. Since traditional devices usually only support simple button operations or touch controls, it is difficult for users to achieve complex interaction needs, especially in scenarios where image and voice input need to be processed simultaneously. This application proposes a multimodal interaction solution by integrating a sound acquisition module and combining it with an intelligent voice wake-up and intent recognition mechanism. During specific implementation, the device will automatically cache the collected target image and monitor whether the voice input contains a preset wake-up word. When the wake-up word is detected, the voice will be converted into text and the user's intention will be analyzed. The image and text content will be intelligently combined according to the intention. It can also actively stop detection and prompt the user when the voice input times out. The intelligent voice control and content combination mechanism improves the flexibility of use and realizes the coordinated processing of images and voice. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a schematic diagram of the overall system architecture of a portable cognitive device based on artificial intelligence according to an embodiment of the present application; Figure 2 This is a flowchart of intelligent page type recognition of a portable cognitive device based on artificial intelligence in an embodiment of the present application; Figure 3This is a flowchart of a graphic and text combination processing of a portable cognitive device based on artificial intelligence in an embodiment of the present application; Figure 4 This is a flowchart of a QR code information extraction process of a portable cognitive device based on artificial intelligence according to an embodiment of the present application; Figure 5 This is a flowchart of reading habit modeling and prediction of a portable cognitive device based on artificial intelligence in an embodiment of the present application; Figure 6 This is a flow chart of environmental adaptability adjustment of an artificial intelligence-based portable cognitive device according to an embodiment of the present application; Figure 7 This is a flowchart of a personalized reading difficulty assessment of a portable cognitive device based on artificial intelligence according to an embodiment of the present application; Figure 8 This is a flowchart of user proficiency adaptation of a portable cognitive device based on artificial intelligence according to an embodiment of the present application; Figure 9 This is a flowchart of content cognitive complexity analysis of a portable cognitive device based on artificial intelligence in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "said," "above," "the," and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to any or all possible combinations comprising one or more of the listed items.

[0026] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0027] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.

[0028] In the first aspect, the present application provides a portable cognitive device based on artificial intelligence, referring to Figure 1The cognitive device includes a shell, an image acquisition module, a light projection module, a processor, a sound playback module and a communication module, wherein the image acquisition module is used to acquire a target image, and the light projection module is used to project an optical positioning mark indicating a framing area into the field of view of the image acquisition module, and the optical positioning mark is used to indicate the effective acquisition area of ​​the image acquisition module. The processor is used to control the image acquisition module to acquire a target image for processing in response to a trigger signal, and can also send the target image and / or processing result to the artificial intelligence service for processing through the communication module, and output the processing result through the sound playback module. It is understandable that the optical positioning mark includes but is not limited to a dot cursor, a line cursor, a box cursor, a corner cursor or a combination thereof.

[0029] Furthermore, the device also includes a power management module arranged in the shell, which is used to control the power supply of various parts of the system and the battery charging management to ensure stable operation of the device. In terms of triggering methods, this embodiment provides a variety of optional interactive schemes: the user can activate the light projection module to display the framing light frame by lightly pressing the touch switch, and press the touch switch again to capture the image after adjusting the position; or activate the framing light frame by pressing and holding the touch switch, and release the touch switch to trigger image capture after adjusting the position. The light projection module automatically turns off the projection of the optical positioning mark after the image capture is completed to reduce energy consumption. The processor realizes precise control of the light projection module through the power management module, and only turns on the projection during the necessary framing stage.

[0030] In one embodiment, referring to Figure 2 , the processor is also used to: S210: Detect the position of the optical positioning mark in the target image and determine the effective acquisition area.

[0031] S220: When the position of the optical positioning mark cannot be detected, determine whether the distance between the device and the target is too close, and provide a prompt through the sound playback module.

[0032] S230: Detect text areas within the effective acquisition area; identify page types based on the distribution of the text areas, where the page types include cover pages, text pages, or other pages.

[0033] When recognizing text areas, the processor automatically detects and excludes incomplete text areas at the edges of the image, ensuring that only complete and readable text content is processed. For text areas that span pages or are partially obscured, the system uses a completeness detection algorithm to filter them out, avoiding the effect of reading out of context.

[0034] Specifically, the intelligent reading process executed by the processor includes two major links: image processing and voice output. In the image processing link, the device first obtains the target image through the image acquisition module and pre-processes the image, including basic operations such as size scaling and color space conversion. Subsequently, the processor detects the position of the optical positioning mark in the target image to determine the effective acquisition area. When the optical positioning mark cannot be detected, the system determines that the distance between the device and the target may be too close. At this time, a prompt sound is issued through the sound playback module to guide the user to adjust the position. After determining the effective acquisition area, the system performs text area detection and intelligent error correction in the area. Intelligent error correction corrects possible text recognition errors through context understanding and semantic analysis. Especially for content with poor printing quality or special fonts, the large model can perform accurate error correction processing in combination with the context. Text area detection identifies the page type based on the distribution characteristics of the text area and classifies it as a cover page, text page or other page type. For the cover page, the system performs error correction and extracts the main and subtitle information. For the body page, the system performs sentence segmentation after error correction and uses a context-cohesion algorithm to remove any overlap with previously read content. For other types of pages (such as the copyright page and back cover), the system uses a voice playback module to prompt the user that they no longer need to read. Finally, the system converts the processed text content into voice data and outputs it through the voice playback module. The entire processing process is highly automated, effectively identifying and processing different types of page content, providing users with a smooth reading experience.

[0035] Furthermore, the system adopts a sliding window strategy to detect and track the center position of the cursor in multiple consecutive frames. Specifically, the system maintains a time window of length N, and records the coordinates of the center of the light frame (x_i, y_i) detected in each frame of the window. For the newly detected position, the Euclidean distance between it and each position in the window is first calculated, and a threshold T is set to judge the position mutation. When the distance exceeds the threshold T, the system determines the motion type by calculating the acceleration characteristics of the position, and performs a second-order difference on the position difference sequence of adjacent frames in the window to obtain an acceleration sequence. If the acceleration changes gently, it is judged as active movement of the user; if the acceleration changes violently, it is judged as jitter. For situations determined to be jitter, the system performs position correction based on the exponentially weighted moving average algorithm, that is, the new cursor center position P_t=α*P_measured+(1-α)*P_t-1, where P_t represents the coordinates of the cursor center position after smoothing of the t-th frame; P_measured represents the coordinates of the cursor center position actually measured in the t-th frame; P_t-1 represents the coordinates of the cursor center position after smoothing of the t-1st frame; α is the smoothing coefficient, and its value range is [0,1]. When severe jitter is detected, α takes a smaller value (such as 0.2) to enhance the smoothing effect. When slight jitter is detected, α takes a larger value (such as 0.8) to maintain response sensitivity.

[0036] In one embodiment, referring to Figure 1 The device also includes a sound collection module arranged in the housing, referring to Figure 3 , the processor is also used to: S310 , buffering the target image acquired by the image acquisition module.

[0037] S320: Detect whether the voice input of the sound collection module contains a preset wake-up word.

[0038] S330: When a wake-up word is detected, convert the voice input into text content and analyze the text content to determine the user's intention.

[0039] S340: Selectively combine the cached target image with the text content according to the user's intention.

[0040] S350: When the input of the sound collection module exceeds the preset time, the detection is stopped and a prompt is given through the sound playing module.

[0041] Specifically, the processor executes the following cognitive learning process: First, the user flicks a switch to activate the light projection module, which projects an optical positioning marker. When the user aligns a learning target (such as an image, physical object, or exercise) with the light frame, the processor controls the image acquisition module to capture and cache the target image. Next, the user can ask a question via voice (such as "What is this?" or "How do I solve this problem?"). The sound acquisition module monitors the voice input in real time. After detecting the preset wake-up word, the processor converts the voice input into text and analyzes the user's intent. Based on the analysis, the processor selectively combines the cached target image with the question text to form a complete interaction request. For example, if a user points to an image of an umbrella and asks "What is this?", the system combines the image with the question and submits it to the intelligent processing module, generating a response such as "This is an umbrella." The user can continue to ask, "What is an umbrella used for?" The system will continue to respond and output the answer through the sound playback module. To optimize the interactive experience, if the sound acquisition module's input exceeds a preset time (for example, 30 seconds without a new question), the system stops detecting and notifies the user through the sound playback module that the conversation has ended, while also turning off the projected light frame. This interactive mode is particularly suitable for children's cognitive learning scenarios. It helps users acquire knowledge through natural dialogue, and realizes the intelligence and humanity of human-computer interaction.

[0042] Furthermore, the sound playback module supports multiple tone configurations, including user-defined tones, personalized tone cloning based on user audio samples, etc. Users can choose or customize the required tone according to actual needs. For example, it is particularly suitable for parents to customize exclusive reading tones for their children.

[0043] In one embodiment, referring to Figure 4 , the processor is also used to: S410 , preprocessing the target image, including image scaling and color space conversion.

[0044] S420: Recognize the QR code content in the target image.

[0045] S430. If the QR code content is recognized, the QR code content is sent to the artificial intelligence service to determine whether the URL in the QR code is safe and compliant, filter illegal URLs, and analyze and filter the content pointed to by the QR code.

[0046] S440: Extract playable content from the processing result of the artificial intelligence service.

[0047] S450: When illegal content is detected, a warning is issued through the sound playing module.

[0048] Specifically, after a user aligns the optical positioning markers of the light projection module with the target QR code, the processor first pre-processes the target image captured by the image acquisition module, including scaling it to an appropriate size and performing color space conversion to improve subsequent recognition accuracy. The pre-processed image then enters the QR code recognition phase. Once the QR code content is recognized, the processor immediately sends it to the AI ​​service for security verification. Within the AI ​​service, the system first determines whether the URL contained in the QR code is safe and compliant, filtering out illegal URLs that may pose security risks. For compliant URLs, the system further analyzes and filters the specific content they point to to ensure its safety and appropriateness. For example, when a user scans a QR code next to a museum exhibit, the system verifies the legitimacy of the link and extracts relevant information about the exhibit. If any illegal or inappropriate content is detected during processing, the processor immediately issues a warning through the audio playback module. For content that passes security verification, the system extracts the playable portion, such as the exhibit's historical background and creation story, converts it into audio, and then outputs it through the audio playback module. This intelligent filtering mechanism significantly improves the device's user safety, allowing users to confidently access supplementary information about various venues.

[0049] In one embodiment, referring to Figure 5 , the processor is also used to: S510: Establish a user reading behavior model and record reading habit data including device movement frequency, average reading time, and page turning interval.

[0050] S520: Predict the type and reading difficulty of the next page based on the reading habit data.

[0051] S530: Based on the prediction result, adaptively adjust the display parameters of the optical positioning mark and the voice playback speed.

[0052] Specifically, the processor continuously collects and analyzes the user's reading behavior data to establish a personalized reading behavior model. This model records multi-dimensional reading habit data, including device movement frequency, average reading time, page turning intervals, etc. Based on this historical data, the processor can intelligently predict the type of the next page (such as whether it is the beginning of a new chapter, whether it contains complex illustrations, etc.) and the reading difficulty. The processor then automatically adjusts the device parameters based on the prediction results: when it is predicted that the next page is more difficult, the system will appropriately reduce the voice playback speed to give the user more time to understand; when it detects that the user is reading in a dark environment, it will automatically increase the brightness of the optical positioning mark; when it predicts that the user is about to enter a long reading, it will optimize the light frame size to reduce eye fatigue.

[0053] In one embodiment, referring to Figure 6 , the processor is also used to: S610: Detect ambient light intensity and ambient noise level.

[0054] S620: Evaluate the suitability of the current reading environment based on the reading habit data.

[0055] S630: When the environmental suitability is lower than a preset threshold, adjust the brightness of the optical positioning mark and the volume of the sound playback module to help the user maintain attention.

[0056] Specifically, the system uses sensors to detect ambient light intensity and noise levels in real time. These parameters are then combined with data on the user's reading habits to comprehensively assess the suitability of the current reading environment. This assessment utilizes a multi-parameter weighted scoring mechanism. For example, light intensity can be divided into multiple ranges and assigned different scores, which are then multiplied by the corresponding weights to create a composite score. When the system detects that a user is reading in low light, it calculates an environmental suitability score. When this score falls below a preset threshold, the processor automatically increases the brightness of the optical positioning markers based on the score difference according to a pre-set algorithm, enhancing the visibility of the viewing area. Similarly, in noisy environments, the system dynamically adjusts the volume gain based on a preset noise floor threshold, using methods such as logarithmic scaling, and sets appropriate volume limits. This intelligently adjusts the volume of the audio playback module to ensure clear audio content. This adaptive environmental mechanism not only helps users maintain focus in a variety of environments but also reduces visual and auditory fatigue, making it particularly suitable for extended reading sessions. By dynamically adjusting device parameters, the system effectively addresses the inability of traditional reading aids to cope with complex environmental changes.

[0057] In one embodiment, referring to Figure 7 , the processor is also used to: S710: Identify the user's reading proficiency based on the reading habit data.

[0058] S720, dynamically adjust the optimal acquisition distance range between the device and the target, the tracking sensitivity of the optical positioning marker, and the interactive frequency of voice playback based on reading proficiency.

[0059] Specifically, the system identifies the user's reading proficiency level by analyzing the user's reading habit data, including dimensions such as device operation stability, alignment speed, and interactive response time. For example, for first-time users, the system detects that the device may not be held stably enough. At this time, it will automatically relax the optimal acquisition distance range between the device and the target, and at the same time reduce the tracking sensitivity of the optical positioning marker to provide a larger operational tolerance space; for experienced users, the system will narrow the effective acquisition distance range and increase the sensitivity of the light frame tracking to achieve faster and more accurate framing. In addition, the system will also adjust the interactive frequency of voice playback according to the user's proficiency: provide more operation prompts and guidance for novice users, and reduce unnecessary interactive prompts for experienced users to make the reading process smoother. This dynamic adjustment mechanism based on proficiency effectively improves the applicability of the device and the user experience.

[0060] In one embodiment, referring to Figure 8 , the processor is also used to: S810: Statistically analyze the duration of users' stay on different types of pages and the number of times they repeatedly read the pages, and obtain statistical analysis results.

[0061] S820. Establish a personalized reading difficulty assessment model based on the statistical analysis results.

[0062] S830: Based on the reading difficulty assessment model, adjust the dwell time of the optical positioning mark and optimize the speech speed and number of repetitions of the voice playback.

[0063] Specifically, the system continuously records and analyzes users' reading behavior on different types of pages, including key indicators such as dwell time and number of repeated readings. For example, the system may find that users spend an average of a longer time dwelling on science and technology content, but fewer times rereading story content. Based on these statistical analysis results, the processor builds a personalized reading difficulty assessment model. Furthermore, the system dynamically adjusts device parameters according to this assessment model: for content types that users often need to read repeatedly, the system will automatically extend the dwell time of the optical positioning mark, reduce the voice playback speed, and appropriately increase the number of repeated playbacks; for content that users can easily understand, a faster speech speed and shorter dwell time are used to improve reading efficiency.

[0064] In one embodiment, referring to Figure 9, the optical positioning mark is a line cursor, a frame cursor, or a combination thereof; the processor is further configured to: S910: Detect the layout features of the text area, divide the optical positioning mark into multiple sub-areas, and adaptively adjust the size of each sub-area according to the cognitive complexity of the text content.

[0065] S920: For the content in each sub-region, identify the unfamiliarity and grammatical structure of the characters and determine the corresponding cognitive difficulty.

[0066] S930. Adjust the dwell time and reading speed of each sub-area according to the cognitive difficulty.

[0067] Specifically, the system first analyzes the typesetting characteristics of the text area in the target image, such as paragraph layout, font size variations, and mixed text and image layout. Based on this, the optical positioning markers are intelligently divided into multiple sub-areas. The size of each sub-area is not simply divided equally, but is dynamically adjusted based on the cognitive complexity of the text content. For example, when a text is detected to contain both chart captions and text descriptions, the system will divide it into sub-areas of different sizes accordingly. The system then conducts an in-depth analysis of the text content within each sub-area, identifying features such as the frequency of occurrence of uncommon characters and the complexity of sentence structure, and assigning different cognitive difficulty levels to each sub-area. Based on these difficulty assessment results, the system further optimizes the reading experience: for sub-areas containing a large number of uncommon characters, the light frame dwell time is appropriately extended and the voice playback speed is reduced; for simple descriptive text, a faster playback rhythm is adopted. This refined processing mechanism based on content characteristics enables the device to more intelligently guide the user's reading rhythm, providing a reading experience that is more in line with cognitive laws.

[0068] It's understandable that, depending on the device's hardware configuration, the processor can selectively execute some algorithms on the device, including basic image preprocessing, optical marker detection, preliminary text area recognition, and simple voice processing tasks. Given sufficient computing power on the device, the proper allocation of computing tasks between local and cloud computing can significantly improve system responsiveness and offline usability.

[0069] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0070] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A portable cognitive device based on artificial intelligence, characterized in that: include: case; An image acquisition module, disposed in the housing, for acquiring a target image; a light projection module, disposed in the housing, for projecting an optical positioning mark indicating a viewing area into the field of view of the image acquisition module; a processor, disposed in the housing; A sound playing module is disposed in the housing; a communication module, disposed in the housing; Among them, the optical positioning mark is used to indicate the effective acquisition area of ​​the image acquisition module, the processor is used to control the image acquisition module to acquire the target image and process it, and send the processing results to the server through the communication module for subsequent processing.

2. The portable cognitive device based on artificial intelligence according to claim 1, characterized in that The processor is further configured to: Detecting the position of the optical positioning mark in the target image to determine the effective acquisition area; When the position of the optical positioning mark cannot be detected, determining whether the distance between the device and the target is too close, and providing a prompt through the sound playing module; Detecting a text area within the effective acquisition area; Based on the distribution position of the text area, the page type is identified, and the page type includes a cover page, a text page or other pages.

3. The portable cognitive device based on artificial intelligence according to claim 1, characterized in that Also includes: A sound collection module is provided in the housing; The processor is further configured to: caching the target image captured by the image acquisition module; Detecting whether the voice input of the sound collection module contains a preset wake-up word; When the wake-up word is detected, converting the voice input into text content and analyzing the text content to determine user intent; selectively combining the cached target image with the text content according to the user intention; When the input of the sound collection module exceeds a preset time, the detection is stopped and a prompt is given through the sound playing module.

4. The portable cognitive device based on artificial intelligence according to claim 1, characterized in that The processor is further configured to: Preprocessing the target image, wherein the preprocessing includes image scaling and color space conversion; Recognizing QR code content in the target image; If a QR code is identified, it is sent to an artificial intelligence service to determine whether the URL in the QR code is safe and compliant, filter illegal URLs, and analyze and filter the content pointed to by the QR code. extracting playable content from the processing results of the artificial intelligence service; When illegal content is detected, a warning is issued through the sound playing module.

5. The portable cognitive device based on artificial intelligence according to claim 2, characterized in that: The processor is further configured to: Establish a user reading behavior model and record reading habit data including device movement frequency, average reading time, and page turning interval; Predicting the type and reading difficulty of the next page based on the reading habit data; Based on the prediction results, the display parameters of the optical positioning mark and the voice playback speed are adaptively adjusted.

6. The portable cognitive device based on artificial intelligence according to claim 5, characterized in that: The processor is further configured to: Detect ambient light intensity and ambient noise level; evaluating the suitability of the current reading environment based on the reading habit data; When the environmental suitability is lower than a preset threshold, the brightness of the optical positioning mark and the volume of the sound playing module are adjusted to help the user maintain attention.

7. The portable cognitive device based on artificial intelligence according to claim 5, characterized in that: The processor is further configured to: identifying the user's reading proficiency based on the reading habit data; The optimal acquisition distance range between the device and the target, the tracking sensitivity of the optical positioning marker, and the interactive frequency of the voice playback are dynamically adjusted according to the reading proficiency.

8. The portable cognitive device based on artificial intelligence according to claim 5, characterized in that: The processor is further configured to: Statistically analyze the length of time users stay on different types of pages and the number of times they read them repeatedly to obtain statistical analysis results; establishing a personalized reading difficulty assessment model based on the statistical analysis results; Based on the reading difficulty assessment model, the dwell time of the optical positioning mark is adjusted and the speech speed and number of repetitions of the voice playback are optimized.

9. The portable cognitive device based on artificial intelligence according to claim 5, characterized in that: The processor is further configured to: Detecting the layout features of the text area, dividing the optical positioning mark into multiple sub-areas, and adaptively adjusting the size of each sub-area according to the cognitive complexity of the text content; For the content in each sub-area, identify the unfamiliarity and grammatical structure of the text and determine the corresponding cognitive difficulty; The dwell time and reading speed of each sub-area are adjusted according to the cognitive difficulty.