Reading equipment and picture book reading method

By generating updated picture book page text in the reading device and converting it to audio, the problem of picture book text recognition caused by image blurring is solved, enabling efficient picture book reading in poor shooting environments.

CN121919166APending Publication Date: 2026-04-24HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2024-10-22
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing reading devices produce blurry images in poor shooting environments, making it impossible to properly recognize picture book text and affecting picture book reading efficiency.

Method used

Images of picture book pages are acquired using an image acquisition device. Combined with a local collection of electronic picture books and the text of already read picture book pages, updated text of the picture book pages is generated. Audio conversion is then performed, and an audio clip matching the picture book pages is output. Prompts are used to improve the accuracy of text recognition.

Benefits of technology

Even with blurry images, accurately identifying the text on the inner pages of picture books improves reading efficiency and ensures the accuracy and coherence of the story content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121919166A_ABST
    Figure CN121919166A_ABST
Patent Text Reader

Abstract

The invention relates to reading equipment and a picture book reading method. Comprising the following steps: acquiring a picture book inside page image obtained by performing image acquisition on a current picture book inside page which is turned to at present in a picture book by a reading device; under the condition that a local electronic picture book matched with the picture book name of the picture book exists in an electronic picture book set locally stored in the reading equipment, generating a picture book inside page text of a current picture book inside page based on the picture book inside page image, and determining a target picture book inside page with the highest matching degree with the picture book inside page text from the local electronic picture book; under the condition that the matching degree of the picture book inner page text and the target picture book inner page does not reach a matching degree threshold value, generating an updated picture book inner page text of the current picture book inner page based on the picture book inner page text and the picture book inner page text of at least one read picture book inner page of the picture book, and performing audio conversion on the updated picture book inner page text to obtain a picture book inner page text; obtaining an audio clip matched with the current picture book inside page; and outputting the audio clip matched with the current picture book inner page. The picture book reading efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of reading device technology, and in particular to a reading device and a picture book reading method. Background Technology

[0002] Picture books are a type of book that primarily uses illustrations and includes a small amount of text. Picture books can not only be used for storytelling and learning, but also comprehensively help children build their character and cultivate multiple intelligences. Using reading devices can further enhance children's experience with picture books.

[0003] Currently, reading devices can capture images of picture books, then use OCR technology to recognize the text in the images, and finally play the corresponding audio.

[0004] However, when the shooting environment is poor and the captured image is blurry, the reading device may be unable to recognize the text properly, thus failing to play the audio and affecting the efficiency of picture book reading. Summary of the Invention

[0005] This application provides a reading device and a picture book reading method that can improve the efficiency of picture book reading.

[0006] On one hand, some embodiments provide a reading device including: an image capture unit, a controller, and an audio output unit. The image acquisition device is configured to acquire an image of the currently turned page of the picture book after the reading device enters the picture book reading process, thereby obtaining a picture book page image. The controller is configured to: if a local electronic picture book with a matching picture book name exists in the locally stored electronic picture book collection, generate the picture book page text of the current picture book page based on the picture book page image, and determine the target picture book page with the highest matching degree from the local electronic picture books; if the matching degree between the picture book page text and the target picture book page does not reach a matching degree threshold, generate an updated picture book page text of the current picture book page based on the picture book page text and the picture book page text of at least one read picture book page of the picture book, and perform audio conversion on the updated picture book page text to obtain an audio segment matching the current picture book page; the audio output device is configured to output the audio segment matching the current picture book page.

[0007] In this embodiment, if a local electronic picture book with a matching picture book name exists in the locally stored collection of electronic picture books, the picture book inner page text of the current picture book inner page is generated based on the inner page image. The target picture book inner page with the highest matching degree with the inner page text is determined from the local electronic picture books. If the matching degree between the inner page text and the target inner page does not reach the matching degree threshold, an updated inner page text of the current picture book inner page is generated based on the inner page text and the inner page text of at least one read inner page of the picture book. The updated inner page text is then converted into audio to obtain an audio segment that matches the current inner page. This allows for more accurate inner page text even when the inner page image is blurry, facilitating smooth reading and further improving picture book reading efficiency.

[0008] In some embodiments, the step of generating updated picture book page text of the current picture book page based on the picture book page text and the picture book page text of at least one read picture book page of the picture book is further configured to: generate a first prompt, the first prompt being used to instruct the generation of updated picture book page text of the current picture book page using the picture book page text and the picture book page text of at least one read picture book page of the picture book; and generate updated picture book page text of the current picture book page based on the first prompt, the picture book page text, and the picture book page text of at least one read picture book page of the picture book.

[0009] In this embodiment, by using the first prompt and the text of at least one read page of the picture book, the content of the next page can be identified based on the previous page of the picture book's storyline. This ensures that the correct story content is obtained even when the data collection is inaccurate (blurry images, incomplete shooting angles), or the story can be expanded.

[0010] In some embodiments, the step of generating the picture book inner page text of the current picture book inner page based on the picture book inner page image is further configured to: generate a second prompt, the second prompt being used to indicate the recognition of text in the picture book inner page image; and generate the picture book inner page text of the current picture book inner page based on the second prompt and the picture book inner page image.

[0011] In this embodiment, the picture book inner page text of the current picture book inner page is generated based on the second prompt and the picture book inner page image, thereby improving the accuracy of the generated picture book inner page text by utilizing the second prompt.

[0012] In some embodiments, determining the target picture book page with the highest matching degree with the picture book page text from the local electronic picture book is further configured to: split the picture book page text to obtain multiple text fragments; and determine the target picture book page with the highest matching degree with the picture book page text from the local electronic picture book based on the multiple text fragments.

[0013] In this embodiment, the target picture book page that matches the text of the picture book page is determined from the local electronic picture book based on multiple text fragments, thereby improving the accuracy of the target picture book page.

[0014] In some embodiments, the reading device is further configured to: when the matching degree between the text on the inner page of the picture book and the target inner page of the picture book reaches a matching degree threshold, determine an audio segment pre-generated for the target inner page of the picture book, and obtain an audio segment that matches the current inner page of the picture book.

[0015] In this embodiment, when the matching degree between the text on the inner page of the picture book and the target inner page of the picture book reaches the matching degree threshold, it means that a target inner page of the picture book that matches the current inner page of the picture book has been found. Thus, the audio segment pre-generated for the target inner page of the picture book is determined, and an audio segment matching the current inner page of the picture book is obtained. This not only improves the accuracy of the obtained audio segment, but also enables the rapid reading of the picture book using the local electronic picture book, thereby improving the efficiency of picture book reading.

[0016] In some embodiments, the output of the audio segment matching the current picture book page is further configured to: output the audio segment matching the current picture book page when the current reading state of the target picture book page is unread; the controller is further configured to: update the reading state of the target picture book page from unread to reading when outputting the audio segment.

[0017] In this embodiment, if the current reading status of the target picture book page is "not read", an audio clip can be output, thereby avoiding repeatedly reading the picture book page that has already been read. When the audio clip is output, i.e. when reading begins, the reading status is updated to "reading", so that the reading status can be updated in a timely manner.

[0018] In some embodiments, the controller is further configured to: output page-turning prompts when the current reading status of the target picture book page is "read"; and exit the picture book reading process if page-turning prompts are output multiple times consecutively.

[0019] In this embodiment, if the current reading status of the target picture book page is "read," it means that the current page has been read, and a translation prompt message is output to promptly remind the user to turn the page. Furthermore, if the page-turning prompt message is output multiple times consecutively, it may indicate that the user is no longer engaged in reading and has exited the picture book reading process. This avoids unnecessary waste of computer resources caused by the reading device remaining continuously in the picture book reading process.

[0020] In some embodiments, before capturing an image of the currently turned page in the picture book, the image capture device is further configured to capture an image of the picture book cover to obtain a picture book cover image; the controller is further configured to perform text recognition on the picture book cover image to obtain the picture book name.

[0021] In this embodiment, before capturing an image of the currently turned page in the picture book, the picture book name is identified first, which helps to determine the text and audio clips on the picture book page later.

[0022] In some embodiments, the step of performing text recognition on the picture book cover image to obtain the picture book name is further configured to: perform text recognition on the picture book cover image to obtain cover recognition text; determine a plurality of preset texts, each of the preset texts being used to characterize an inaccurate result not identified; and determine the picture book name based on the cover recognition text and the plurality of preset texts.

[0023] In this embodiment, since the preset text is used to represent the inaccurate result, the cover recognition text can be accurately determined as the picture book title based on the preset text, thus improving the accuracy of the picture book title.

[0024] In some embodiments, determining the picture book name based on the cover recognition text and the plurality of preset texts is further configured to: determine the character length of the cover recognition text and determine a preset character length threshold; and determine the picture book name based on the cover recognition text, the character length, the character length threshold, and the plurality of preset texts.

[0025] In this embodiment, the picture book title is determined by combining the character length and preset text, which improves the accuracy of the picture book title.

[0026] In some embodiments, the controller is further configured to: if no local electronic picture book matching the picture book name exists in the locally stored collection of electronic picture books, generate the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book, and perform audio conversion on the picture book inner page text to obtain an audio segment matching the current picture book inner page.

[0027] This embodiment integrates online picture book reading and local electronic picture book reading. After recognizing the picture book cover and obtaining the story title, different modes are selected based on whether a local pre-made picture book is available. Online picture book reading offers greater scalability, especially when the recognized image content is incomplete, leveraging the context learning ability of the image understanding model to quickly expand the story and ensure the enjoyment of reading picture books.

[0028] On the other hand, some embodiments also provide a picture book reading method, applied to a reading device or server. The method includes: after the reading device enters the picture book reading process, acquiring a picture book page image obtained by the reading device from the currently turned picture book page; if there is a local electronic picture book in the local electronic picture book collection stored by the reading device that matches the picture book name of the picture book, generating the picture book page text of the current picture book page based on the picture book page image, and determining the target picture book page with the highest matching degree from the local electronic picture books; if the matching degree between the picture book page text and the target picture book page does not reach the matching degree threshold, generating an updated picture book page text of the current picture book page based on the picture book page text and the picture book page text of at least one read picture book page of the picture book, and performing audio conversion on the updated picture book page text to obtain an audio segment matching the current picture book page; and outputting the audio segment matching the current picture book page.

[0029] In this embodiment, if a local electronic picture book with a matching picture book name exists in the collection of electronic picture books stored locally on the reading device, the inner page text of the current picture book is generated based on the inner page image. The target picture book inner page with the highest matching degree with the inner page text is determined from the local electronic picture books. If the matching degree between the inner page text and the target inner page does not reach the matching degree threshold, an updated inner page text of the current picture book inner page is generated based on the inner page text and the inner page text of at least one read inner page of the picture book. The updated inner page text is then converted into audio to obtain an audio segment that matches the current inner page. This allows for more accurate inner page text even when the inner page image is blurry, facilitating smooth reading and further improving picture book reading efficiency.

[0030] On the other hand, some embodiments provide a reading device, including: an image acquisition unit configured to acquire an image of the currently turned page of a picture book after the reading device enters the picture book reading process, thereby obtaining a picture book page image; a controller configured to: generate picture book page text of the current picture book page based on the picture book page image and picture book page text of at least one read picture book page of the picture book if no local electronic picture book matching the picture book name exists in a locally stored collection of electronic picture books, and perform audio conversion on the picture book page text to obtain an audio segment matching the current picture book page; and an audio output unit configured to output the audio segment matching the current picture book page.

[0031] In this embodiment, if no locally stored electronic picture book with a matching picture book name exists in the locally stored collection of electronic picture books, the text of the current picture book page is generated based on the picture book page image and the text of at least one read picture book page. The text is then converted into audio to obtain an audio segment that matches the current picture book page. This allows for the generation of the current picture book page text based on already read picture book pages, even when the picture book being read is not stored locally, thus ensuring a reasonable picture book content for smooth reading and improving reading efficiency.

[0032] In some embodiments, the step of generating the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book is further configured to: generate a third prompt, the third prompt being used to instruct the generation of the picture book inner page text of the current picture book inner page using the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book; and generate the picture book inner page text of the current picture book inner page based on the third prompt, the picture book inner page image, and the picture book inner page text of at least one read picture book inner page of the picture book.

[0033] In this embodiment, by using the third prompt and the text of at least one read page of the picture book, the content of the next page can be identified based on the previous page of the picture book's storyline. This ensures that the correct story content is obtained even when the data collection is inaccurate (blurry images, incomplete shooting angles), or the story can be expanded.

[0034] In some embodiments, before capturing an image of the currently turned page in the picture book, the image capture device is further configured to capture an image of the picture book cover to obtain a picture book cover image; the controller is further configured to perform text recognition on the picture book cover image to obtain the picture book name.

[0035] In this embodiment, before capturing an image of the currently turned page in the picture book, the picture book name is identified first, which helps to determine the text and audio clips on the picture book page later.

[0036] On the other hand, some embodiments also provide a picture book reading method, applied to a reading device or server. The method includes: after the reading device enters the picture book reading process, acquiring a picture book page image obtained by the reading device from the currently turned picture book page; if there is no local electronic picture book in the electronic picture book collection stored locally by the reading device that matches the picture book name, generating the picture book page text of the current picture book page based on the picture book page image and the picture book page text of at least one read picture book page of the picture book, and performing audio conversion on the picture book page text to obtain an audio segment matching the current picture book page; and outputting the audio segment matching the current picture book page.

[0037] In this embodiment, if no locally stored electronic picture book with a matching picture book name exists in the locally stored collection of electronic picture books, the text of the current picture book page is generated based on the picture book page image and the text of at least one read picture book page. The text is then converted into audio to obtain an audio segment that matches the current picture book page. This allows for the generation of the current picture book page text based on already read picture book pages, even when the picture book being read is not stored locally, thus ensuring a reasonable picture book content for smooth reading and improving reading efficiency. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1A A schematic diagram illustrating the interaction between a reading device and a user, provided in some embodiments of this application;

[0040] Figure 1B This is a schematic diagram illustrating an operational scenario between a reading device and a control device provided in some embodiments of this application;

[0041] Figure 2 This is a schematic diagram of the hardware configuration of a reading device provided in some embodiments of this application;

[0042] Figure 3This is a schematic diagram of the hardware configuration of the control device provided in some embodiments of this application;

[0043] Figure 4 This is a schematic diagram of the software configuration of a reading device provided in some embodiments of this application;

[0044] Figure 5 A flowchart illustrating a picture book reading method provided in some embodiments of this application;

[0045] Figure 6 A schematic diagram illustrating the principle of a picture book reading method provided in some embodiments of this application;

[0046] Figure 7 System architecture diagram of a picture book reading method provided in some embodiments of this application;

[0047] Figure 8 A flowchart illustrating the process of identifying picture book titles provided for some embodiments of this application;

[0048] Figure 9 A flowchart illustrating a picture book reading method provided in other embodiments of this application;

[0049] Figure 10 A timing diagram illustrating a picture book reading method provided in some embodiments of this application. Detailed Implementation

[0050] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0051] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0052] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0053] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0054] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0055] In this embodiment, the reading device 100 generally refers to a device with image acquisition, image processing, and audio output capabilities. The reading device may also have screen display capabilities. For example, the reading device 100 includes, but is not limited to, reading robots, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, and augmented reality devices.

[0056] Figure 1A This is a schematic diagram illustrating the operation scenarios between the reading device and the user according to some embodiments of this application. The user can directly control the reading device 100. For example, the reading device 100 may provide an operation interface or operation buttons, and the user can directly control or set the reading device 100 on the operation interface.

[0057] Figure 1B This is a schematic diagram illustrating an operational scenario between a reading device and a control device provided in some embodiments of this application. For example... Figure 1B As shown, users can operate the reading device 100 via touch operation, mobile terminal 300, and control device 400. For example, the control device 400 can be a remote control, stylus, or gamepad. Alternatively, users can control the reading device 100 directly without using the mobile terminal 300 and control device 400. For instance, the reading device 100 can provide an interface or operation buttons, allowing users to directly control or configure the reading device 100 on the interface.

[0058] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the reading device 100. It can also function as a communication device for establishing a communication connection with the reading device 100 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the reading device 100 to establish a connection and communicate via network communication protocols, achieving one-to-one control operations and data communication. Furthermore, audio and video content displayed on the mobile terminal 300 can be transmitted to the reading device 100 for synchronized display.

[0059] like Figure 1BThe document also shows that the reading device 100 communicates with the server 200 via various communication methods. This allows the reading device 100 to connect and communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0060] The reading device 100 can provide broadcast television reception function, and can also be equipped with smart network TV function that provides computer support, including but not limited to network TV, smart TV, Internet Protocol TV (IPTV), etc.

[0061] Figure 2 This is a hardware configuration block diagram of a reading device 100 provided in some embodiments of this application.

[0062] In some embodiments, the reading device 100 may include at least one of the following: a tuner 110, a communication device 120, a detector 130, a device interface 140, a controller 150, a display 160, an audio output device, i.e., an audio output device 170, an image acquisition device such as a camera, a memory, a power supply, and a user input interface.

[0063] In some embodiments, detector 130 is used to acquire signals of the external environment or interaction with the outside world. For example, detector 130 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 130 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 130 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0064] In some embodiments, the display 160 includes display function components for presenting images and driving components for driving image display. The display 160 is used to receive and display image signals output from the controller 150. For example, the display 160 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0065] In some embodiments, the communication device 120 is a component used to communicate with external devices or the server 200 according to various communication protocol types. The reading device 100 may have multiple communication devices 120 depending on the supported communication methods. For example, when the reading device 100 supports wireless network communication, it may have a communication device 120 with WiFi functionality. When the reading device 100 supports Bluetooth connection communication, it needs to have a communication device 120 with Bluetooth functionality.

[0066] The communication device 120 enables the reading device 100 to communicate with external devices or the server 200 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the reading device 100 to external devices. Wireless connections utilize wireless signals or wireless networks. The reading device 100 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0067] In some embodiments, the controller 150 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 150 controls the operation of the reading device and responds to user operations through various software control programs stored in memory. The controller 150 controls the overall operation of the reading device 100.

[0068] In some embodiments, the controller 150 and the tuner 110 may be located in different separate devices, that is, the tuner 110 may also be located in an external device of the main device where the controller 150 is located, such as an external set-top box.

[0069] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 160, and the user input interface receives user input commands through the graphical user interface (GUI).

[0070] In some embodiments, the audio output device, i.e., the audio output unit 170, can be the built-in speaker of the reading device 100 or an external audio output device connected to the reading device 100. For the external audio output device connected to the reading device 100, the reading device 100 may also be provided with an external audio output terminal, through which the audio output device can be connected to the reading device 100 to output sound from the reading device 100.

[0071] In some embodiments, the user input interface 180 can be used to receive instructions from user input.

[0072] Figure 3 This is a hardware configuration block diagram of a control device provided in some embodiments of this application. For example... Figure 3 As shown, the control device 400 may include: a controller 410, a communication interface 430, a user input / output interface, a memory, and a power supply.

[0073] The control device 400 is configured to control the reading device 100, and to receive user input operation commands and convert the operation commands into commands that the reading device 100 can recognize and respond to, thus acting as an intermediary for interaction between the user and the reading device 100.

[0074] In some embodiments, the control device 400 may be a smart device. For example, the control device 400 may be equipped with various applications for controlling the reading device 100 according to user needs.

[0075] In some embodiments, such as Figure 1B As shown, a mobile terminal 300 or other smart electronic device can perform similar functions to control device 400 after installing an application for controlling the reading device 100.

[0076] The controller 410 includes a processor 412, RAM 413, ROM 414, a communication interface 430, and a communication bus. The controller 410 is used to control the operation of the control device 400, as well as the communication and cooperation between internal components and the external and internal data processing functions.

[0077] Under the control of the controller 410, the communication interface 430 enables communication of control signals and data signals with the reading device 100. The communication interface 430 may include at least one of other near-field communication modules such as WiFi chip 431, Bluetooth module 432, and NFC module 433.

[0078] User input / output interface 440, wherein the input interface includes at least one of other input interfaces such as microphone 441, touchpad 442, sensor 443, and button 444.

[0079] In some embodiments, the control device 400 includes at least one of a communication interface 430 and an input / output interface 440. The control device 400 is configured with a communication interface 430, such as a WiFi, Bluetooth, or NFC module, which can encode user input commands via WiFi, Bluetooth, or NFC protocols and send them to the reading device 100.

[0080] The memory 490 is used to store various operating programs, data, and applications for driving and controlling the control device 400 under the control of the controller. The memory 490 can also store various control signal commands input by the user.

[0081] Power supply 480 is used to provide operating power support for the various components of control device 400 under the control of the controller.

[0082] In some embodiments, the reading device 100 may run an operating system to enable user interaction. The operating system is a computer program used to manage and control the hardware and software resources of the reading device 100. The operating system can (control the reading device) provide a user interface, allowing users to interact with the reading device 100 and supporting the running of various applications.

[0083] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for reading devices.

[0084] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0085] In some embodiments, the application layer provides services and interfaces for applications, enabling the reading device 100 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0086] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0087] like Figure 4As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0088] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0089] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0090] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the reading device 100. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 4 As shown, hardware drivers can be configured in the kernel layer. The kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0091] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the reading device 100 in this application embodiment. Depending on the functions of the reading device, the type of operating system, and other factors, the number of layers and the specific type of the operating system may take other forms.

[0092] Based on this, in some embodiments, this application provides a picture book reading method. This method can be executed by a reading device or a server, or it can be jointly executed by the reading device and the server. The reading device can be understood as a voice-activated robot, and a picture book reading tool can be installed on the reading device. The picture book reading method provided in this application can be implemented by the reading device through the picture book reading tool, such as... Figure 5 As shown, the method includes:

[0093] Step 502: After the reading device enters the picture book reading process, acquire the picture book page image obtained by the reading device from the image capture of the currently turned picture book page.

[0094] The reading device contains an image acquisition device, such as a camera, which can be used to capture image data of the picture book, providing image data for the image understanding model. The images of the picture book's inner pages are captured by this image acquisition device. The image understanding model can be a neural network model used to understand images; for example, it can extract text from an image through image understanding. The image understanding model can be a large model.

[0095] The inner pages of a picture book refer to all pages in a picture book except for the cover. In step 502, "picture book" refers to a physical picture book in the user's hand, or a picture book displayed on the screen of an electronic device that allows page turning. The picture book reading process refers to the process of reading a picture book.

[0096] Specifically, the reading device can enter the picture book reading process based on the user's voice instructions. For example, if the user says "I want to read a picture book," the reading device will launch the picture book reading application and enter the picture book reading process.

[0097] In some embodiments, the reading device can guide the user via voice to adjust the position of the picture book in their hands, or guide the user to place the picture book in a designated location. For example, the reading device can ask, "Is there a picture book you can read?" If the user replies "Yes," "The picture book is ready," or "Please read the story of XXX," the reading device can analyze the user's response. If the analysis determines that the user has accurately prepared the picture book, the reading process begins.

[0098] In some embodiments, with the image acquisition device activated and displaying normally, the reading device can provide voice prompts to the user to place the picture book in a designated location, such as by saying, "Please place the picture book directly above the robot," where "robot" refers to the reading device. The user then places the picture book in the designated location according to the voice prompts.

[0099] In some embodiments, when the image acquisition device displays an abnormal state, the reading device can provide an abnormal prompt via voice, such as issuing a voice prompt "Please confirm whether the camera is connected correctly".

[0100] In some embodiments, the image acquisition device can acquire images at preset time intervals. For example, if the time interval is 5 seconds (s), it means that the image acquisition device acquires an image every 5 seconds, that is, it performs a camera data scan every 5 seconds.

[0101] Step 504: If there is a local electronic picture book in the collection of electronic picture books stored locally on the reading device that matches the picture book name, generate the picture book inner page text of the current picture book inner page based on the picture book inner page image, and determine the target picture book inner page with the highest matching degree with the picture book inner page text from the local electronic picture books.

[0102] The locally stored collection of electronic picture books is pre-stored on the reading device. This collection can include multiple picture books in electronic format, each with its own corresponding audio file. The audio file stores the audio data for each page of the picture book; by playing the audio data for each page, the text on that page can be read. The reading device locally stores the audio file corresponding to each picture book in the electronic picture book collection.

[0103] The text on the inner pages of a picture book is the recognition or generation of text contained within the images on the inner pages. The higher the accuracy of the recognition or generation, the closer the text on the inner pages is to the text contained within the images, for example, they can be identical. Local electronic picture books refer to picture books in electronic form from a collection of electronic picture books stored locally. A local electronic picture book matching the picture book name is the same picture book as the one currently being read; the difference is that one is an electronic picture book stored on the reading device, and the other is the picture book the user needs to read.

[0104] Specifically, given the picture book title, the reading device can use fuzzy matching to query whether there is a locally stored electronic picture book that matches the title. Alternatively, the reading device can calculate the similarity between the picture book title and the title of each local electronic picture book, and select the local electronic picture books with a similarity score greater than a preset similarity threshold as the matching titles.

[0105] In some embodiments, a local electronic picture book can be represented by the Java class PictureBookBean, which contains:

[0106] {

[0107] The string "title" represents the name of the picture book.

[0108] List <storycapturebean>This indicates the corresponding chapter content or inner page content.

[0109]

[0110] }

[0111] StoryCaptureBean{

[0112] String content; Story content

[0113] String url: Audio file path

[0114] Int status: The status (reading, readed, unRead) records whether the data has been read.

[0115] }

[0116] In some embodiments, the reading device can input the captured images of the picture book's inner pages into an image understanding model, or send the captured images of the picture book's inner pages to a server, which then inputs the captured images into the image understanding model, analyzes the images, and outputs the text for the current picture book page. Alternatively, the reading device itself can have an image understanding model deployed within it, allowing it to directly input the captured images of the picture book's inner pages into the model to obtain the text output by the model.

[0117] In some embodiments, the reading device or server can generate a prompt indicating that the image of the picture book's inner page should be recognized to identify the text within the image, and that if recognition fails, a prompt such as "known" should be output. The reading device or server can then input the prompt and the image of the picture book's inner page into an image understanding model to obtain the text of the picture book's inner page output by the model.

[0118] Of course, reading devices or servers can also use OCR (Optical Character Recognition) technology to extract text from the images of the picture book's inner pages, thus obtaining the text of the picture book's inner pages. OCR technology is a text recognition technology.

[0119] In some embodiments, the reading device or server may determine the matching degree, such as similarity, between the text in each inner page of the local electronic picture book and the text of that inner page, and determine the inner page with the highest matching degree in the local electronic picture book as the target inner page of the picture book.

[0120] Step 506: If the matching degree between the picture book inner page text and the target picture book inner page does not reach the matching degree threshold, based on the picture book inner page text and the picture book inner page text of at least one read picture book inner page, generate the updated picture book inner page text of the current picture book inner page, and perform audio conversion on the updated picture book inner page text to obtain an audio segment that matches the current picture book inner page.

[0121] The matching threshold can be set as needed. Reaching the matching threshold indicates that the text in the target picture book's inner page is consistent with or substantially consistent with the text on the inner page of the picture book; failing to reach the matching threshold indicates that the text in the target picture book's inner page differs significantly from the text on the inner page of the picture book. The matching threshold can be set to, for example, 90%, 96%, or 100%. "Read picture book pages" refers to the picture book pages that have already been read within the currently being read picture book.

[0122] In step 506, the current page in the picture book is not the first page in the picture book. Since there are no already read pages in the picture book, the text for the first page can be generated based on its image. Specifically, the method for generating the text in step 506 can be either using an image understanding model or recognizing it using OCR technology.

[0123] In some embodiments, the reading device may utilize text-to-speech (TTS) technology to convert the text on the pages of a picture book into audio clips. TTS technology is a technology that can convert text information into speech signals.

[0124] In some embodiments, after receiving the updated picture book inner page text, the reading device or server can determine whether there is a picture book inner page text in each read picture book inner page that is consistent with or substantially consistent with the updated picture book inner page text. If not, an audio segment is identified; otherwise, a prompt is made to turn the page.

[0125] Step 508: Output the audio clip that matches the current page of the picture book.

[0126] Specifically, when the server determines the audio segment, it can send the determined audio segment to the reading device, which can then play the audio segment that matches the current page of the picture book, thereby enabling the reading of that page. Conversely, when the reading device determines the audio segment, it can play that audio segment after determining it.

[0127] In some embodiments, after determining the picture book title, the reading device can prompt the user to start reading the picture book, for example, by issuing a voice prompt such as "Start reading the picture book of XXX". Then, for each page of the picture book that the user turns to, the reading device can perform some or all of the steps in steps 502-508 above to realize the reading of the picture book.

[0128] In this embodiment, if a local electronic picture book with a matching picture book name exists in the collection of electronic picture books stored locally on the reading device, the inner page text of the current picture book is generated based on the inner page image. The target picture book inner page with the highest matching degree with the inner page text is determined from the local electronic picture books. If the matching degree between the inner page text and the target inner page does not reach the matching degree threshold, an updated inner page text of the current picture book inner page is generated based on the inner page text and the inner page text of at least one read inner page of the picture book. The updated inner page text is then converted into audio to obtain an audio segment that matches the current inner page. This allows for more accurate inner page text even when the inner page image is blurry, enabling smooth reading and further improving the efficiency and accuracy of picture book reading.

[0129] Based on this, in some embodiments, this application provides a reading device, including an image acquisition unit, a controller, and an audio output unit. The image acquisition unit is configured to acquire an image of the currently turned page of a picture book after the reading device enters the picture book reading process. The controller is configured to: if a local electronic picture book matching the picture book name exists in a locally stored collection of electronic picture books, generate the picture book page text of the current picture book page based on the picture book page image, and determine the target picture book page with the highest matching degree from the local electronic picture books; if the matching degree between the picture book page text and the target picture book page does not reach a matching degree threshold, generate an updated picture book page text of the current picture book page based on the picture book page text and the picture book page text of at least one read picture book page, and perform audio conversion on the updated picture book page text to obtain an audio segment matching the current picture book page; the audio output unit is configured to output the audio segment matching the current picture book page.

[0130] The reading device provided in this application, when a local electronic picture book with a matching picture book name exists in the locally stored collection of electronic picture books, generates the picture book inner page text of the current picture book inner page based on the picture book inner page image. It then identifies the target picture book inner page from the local electronic picture books with the highest matching degree of the inner page text. If the matching degree between the inner page text and the target inner page does not reach a matching degree threshold, it generates an updated picture book inner page text of the current picture book inner page based on the inner page text and the inner page text of at least one read picture book inner page. The updated inner page text is then converted into audio to obtain an audio segment matching the current inner page. This allows for more accurate inner page text even when the inner page image is blurry, facilitating smooth reading and further improving picture book reading efficiency and accuracy.

[0131] In some embodiments, generating updated picture book inner page text for the current picture book inner page based on picture book inner page text and picture book inner page text of at least one read picture book inner page of the picture book is further configured to: generate a first prompt, the first prompt being used to instruct the generation of updated picture book inner page text for the current picture book inner page using picture book inner page text and picture book inner page text of at least one read picture book inner page of the picture book; and generating updated picture book inner page text for the current picture book inner page based on the first prompt, picture book inner page text, and picture book inner page text of at least one read picture book inner page of the picture book.

[0132] The first prompt may be, for example, "You are a children's picture book explanation expert. Please identify the story plot of this page based on the picture book content of the previous chapter and the picture book content of this chapter," or "You are a children's picture book explanation expert. Please identify the story plot of this page based on the picture book content of the previous page and the picture book content of this page, and output the string content. If the recognition fails, return 'known'." A page can be a whole chapter or a part of a chapter.

[0133] Specifically, the reading device or server can input the first prompt, the text of the picture book's inner pages, and the text of at least one read inner page of the picture book into the image understanding model to generate the updated inner page text of the current inner page.

[0134] In this embodiment, by using the first prompt and the text of at least one read page of the picture book, the content of the next page can be identified based on the previous page of the picture book's storyline. This ensures that the correct story content is obtained even when the data collection is inaccurate (blurry images, incomplete shooting angles), or the story can be expanded.

[0135] In some embodiments, generating the picture book inner page text of the current picture book inner page based on the picture book inner page image is further configured to: generate a second prompt, the second prompt being used to indicate the recognition of text in the picture book inner page image; and generate the picture book inner page text of the current picture book inner page based on the second prompt and the picture book inner page image.

[0136] The second prompt could be something like, "You are a children's story narration expert. Accurately understand the following story content and extract the relevant information" or "You are a children's story narration expert. Accurately understand the following story content and extract the relevant information. Output the string content. If recognition fails, return 'known'."

[0137] Specifically, the reading device or server can input the second prompt and the image of the picture book's inner page into the image understanding model to generate the text of the current picture book's inner page.

[0138] In this embodiment, the picture book inner page text of the current picture book inner page is generated based on the second prompt and the picture book inner page image, thereby improving the accuracy of the generated picture book inner page text by utilizing the second prompt.

[0139] In some embodiments, determining the target picture book page with the highest matching degree with the picture book page text from the local electronic picture book is further configured to: split the picture book page text to obtain multiple text fragments; and based on the multiple text fragments, determine the target picture book page with the highest matching degree with the picture book page text from the local electronic picture book.

[0140] The text inside the picture book can be represented as capture_content. For example, the text inside the picture book could be: "One morning, Little Bear woke up in bed as usual, wearing her usual pajamas, and yawned as usual." The page number of the target picture book page in the local e-book and the page number of the currently viewed picture book page are the same, for example, both being page 5.

[0141] Specifically, the reading device can split the text on the inner pages of a picture book into multiple strings based on the character length and punctuation marks, with each string being a text fragment. For example, if an array Ocr_result is used to store the split text fragments, then for example, Ocr_result[0] = One morning, Ocr_result[1] = Little Bear Sister woke up from her bed as usual, ...

[0142] In some embodiments, the reading device can determine the inner pages containing the multiple text fragments obtained from the split text fragments from a local electronic picture book that matches the picture book name, thereby obtaining the target picture book inner pages. For example, Ocr_result is used as a parameter to sequentially select the corresponding List in PictureBookBean{} <storycapturebean>The function checks whether the content is contained within the list; if it is, then the list is determined to be contained within the list. <storycapturebean>The represented picture book page is the target picture book page. Alternatively, for each picture book page in a local electronic picture book that matches the picture book title, the reading device can determine the matching degree between that picture book page and the multiple text fragments. The higher the matching degree, the greater the probability that the picture book page contains the multiple text fragments. The reading device can then identify the picture book page with the highest matching degree as the target picture book page.

[0143] In this embodiment, the target picture book page that matches the text of the picture book page is determined from the local electronic picture book based on multiple text fragments, thereby improving the accuracy of the target picture book page.

[0144] In some embodiments, the reading device is further configured to: when the matching degree between the text on the inner page of the picture book and the inner page of the target picture book reaches a matching degree threshold, determine an audio segment pre-generated for the inner page of the target picture book, and obtain an audio segment that matches the current inner page of the picture book.

[0145] Specifically, each page in the local electronic picture book corresponds to an audio segment. The reading device can read the page by playing the corresponding audio segment.

[0146] In this embodiment, when the matching degree between the text on the inner page of the picture book and the target inner page of the picture book reaches the matching degree threshold, it means that a target inner page of the picture book that matches the current inner page of the picture book has been found. Thus, the audio segment pre-generated for the target inner page of the picture book is determined, and an audio segment matching the current inner page of the picture book is obtained. This not only improves the accuracy of the obtained audio segment, but also enables the rapid reading of the picture book using the local electronic picture book, thereby improving the efficiency of picture book reading.

[0147] In some embodiments, the output of an audio segment matching the current picture book page is further configured to: output an audio segment matching the current picture book page when the current reading state of the target picture book page is unread; the controller is further configured to: update the reading state of the target picture book page from unread to reading when outputting the audio segment.

[0148] The reading device can locally store the reading status of the inner pages of the local electronic picture book. The reading status can be, but is not limited to, unread, reading, or read. For example, the value of status in StoryCaptureBean above represents the reading status. If status = readed, it means that the inner page of the picture book is read; if status = unread, it means that the inner page of the picture book is unread; if status = reading, it means that the inner page of the picture book is reading; and if status = reading, it means that the inner page of the picture book is reading.

[0149] In this embodiment, if the current reading status of the target picture book page is "not read", an audio clip can be output, thereby avoiding repeatedly reading the picture book page that has already been read. When the audio clip is output, i.e. when reading begins, the reading status is updated to "reading", so that the reading status can be updated in a timely manner.

[0150] In some embodiments, the controller is further configured to: output page-turning prompts when the current reading status of the target picture book page is "read"; and exit the picture book reading process if page-turning prompts are output multiple times consecutively.

[0151] The page-turning prompts are used to guide users to turn pages.

[0152] Specifically, the image acquisition device can capture an image of a picture book's inner page at preset time intervals. For each captured image, it can identify the target page and its current reading status. If the target page is already read, it outputs a page-turning prompt and records the number of times the prompt has not been turned (i.e., the number of consecutive page-turning prompts). If the number of times the prompt has not been turned exceeds a preset threshold, the reading process is terminated, and a voice prompt is given to indicate the impending termination, such as saying, "Since you're not reading with me anymore, I'm going to take a break." The threshold can be set as needed, for example, it could be 10 times.

[0153] In this embodiment, if the current reading status of the target picture book page is "read," it means that the current page has been read, and a translation prompt message is output to promptly remind the user to turn the page. Furthermore, if the page-turning prompt message is output multiple times consecutively, it may indicate that the user is no longer engaged in reading and has exited the picture book reading process. This avoids unnecessary waste of computer resources caused by the reading device remaining continuously in the picture book reading process.

[0154] In some embodiments, before capturing an image of the currently turned page in the picture book, the image capture device is also configured to capture an image of the picture book cover to obtain a picture book cover image; the controller is also configured to perform text recognition on the picture book cover image to obtain the picture book name.

[0155] Specifically, the reading device can generate a cover prompt, which instructs the image understanding model to identify the picture book title from the input picture book cover. For example, the prompt could be, "I want you to act as an image recognition expert, recognizing the picture book title from the input picture book cover image," or "Please recognize the picture book title based on the input picture book cover image. The output format should be a string. If recognition fails, return 'known'." The prompt can also define the recognition process; for example, it could include, "Please keep the original text in the image, do not expand the content, and directly output the recognized picture book title." The reading device can input the cover prompt and the picture book cover image into the image understanding model for text recognition. The image understanding model can understand the picture book cover image based on the cover prompt to identify the picture book title. The reading device can then determine the picture book title based on the recognition result output by the image understanding model.

[0156] In this embodiment, before capturing an image of the currently turned page in the picture book, the picture book name is identified first, which helps to determine the text and audio clips on the picture book page later.

[0157] In some embodiments, performing text recognition on the picture book cover image to obtain the picture book name is further configured to: perform text recognition on the picture book cover image to obtain cover recognition text; determine multiple preset texts, each preset text being used to characterize inaccurate results not identified; and determine the picture book name based on the cover recognition text and the multiple preset texts.

[0158] The preset text can be, but is not limited to, "Sorry," "Unrecognizable," "Trying my best to help," or "Cannot see clearly." The cover recognition text is the recognition result output by the image understanding model described above.

[0159] Specifically, if the reading device determines that the cover recognition text contains any preset text, it will determine that the picture book title has not been recognized. For example, if the cover recognition text is "Sorry, we cannot recognize the image content" or "I will try my best to help you recognize it," the reading device will determine that the picture book title has not been recognized. Since the reason for not recognizing the picture book title may be that the captured picture book cover image is unclear, or that the captured picture book cover image is not actually a picture book cover image, the reading device can provide a voice prompt to the user to recapture the picture book cover, for example, by giving a voice prompt such as "Before reading the picture book, I need to look at the cover first."

[0160] In some embodiments, when the cover recognition text does not include any preset text, the reading device can determine that the cover recognition text is the title of the picture book.

[0161] In this embodiment, since the preset text is used to represent the inaccurate result, the cover recognition text can be accurately determined as the picture book title based on the preset text, thus improving the accuracy of the picture book title.

[0162] In some embodiments, determining the picture book name based on the cover recognition text and multiple preset texts is further configured to: determine the character length of the cover recognition text and determine a preset character length threshold; and determine the picture book name based on the cover recognition text, the character length, the character length threshold, and multiple preset texts.

[0163] The character length threshold can be a preset value or determined based on experience; typically, the character length of the picture book title is less than or equal to this threshold. The character length of the cover recognition text refers to the number of characters included in the cover recognition text.

[0164] In some embodiments, when the character length of the cover recognition text is greater than a character length threshold, the reading device determines that the picture book title has not been recognized.

[0165] In some embodiments, if the character length is less than or equal to a preset character length threshold, and none of the preset texts are present in the cover recognition text, then the cover recognition text is determined as the picture book title.

[0166] In some embodiments, such as Figure 8 The diagram illustrates a flowchart for determining the title of a picture book. First, the reading device inputs data such as images (picture book cover image, cover description) into a large model on the server and registers the large model's recognition result on the server. The purpose of registering the large model's recognition result is to ensure that the server returns the large model's recognition result to the reading device. If the large model's recognition result indicates that no picture book title exists, then... Figure 8 If the "success" branch is "no", the server can return a status code, such as "fail", indicating recognition failure to the reading device; otherwise, it can... Figure 8 If the "success" branch is true, the server can return the recognition result of the large model and a status code indicating successful recognition, such as "success," to the reading device. If the reading device receives a status code indicating failed recognition, it will provide a voice prompt to the user. If the reading device receives a recognition result, but the result is not necessarily the picture book title, it may also be unrecognizable. Therefore, the reading device can determine whether there is a result such as "no image text recognized," i.e., whether there is preset text. If so, the reading device will provide a voice prompt to the user. If there is no preset text, it can further determine whether the character length of the recognition result is greater than a character length threshold, such as 15 characters. If so, the reading device will provide a voice prompt to the user; otherwise, the recognition result will be used as the picture book title. The voice prompt to the user may be, for example, "Before reading the picture book, I need to look at the cover."

[0167] In this embodiment, the picture book title is determined by combining the character length and preset text, which improves the accuracy of the picture book title.

[0168] In some embodiments, the controller is further configured to: if no local electronic picture book with a picture book name matching the picture book exists in the locally stored collection of electronic picture books, generate the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book, and perform audio conversion on the picture book inner page text to obtain an audio segment matching the current picture book inner page.

[0169] The method of reading picture books when a matching local e-picture book exists in the locally stored collection can be understood as a local reading method. Conversely, the method of reading picture books when no matching local e-picture book exists in the locally stored collection can be understood as an online reading method. Therefore, it is possible to achieve a reading method that integrates local and online picture book reading.

[0170] Specifically, if there is no local electronic picture book in the locally stored collection that matches the picture book name, the reading device or server can generate the picture book inner page text of the current picture book inner page based on the picture book inner page image, and generate the picture book inner page text of the current picture book inner page based on the picture book inner page text and the picture book inner page text of at least one read picture book inner page of the picture book.

[0171] In this embodiment, even when the currently read picture book is not stored locally, the text of the current picture book page can be generated based on the already read pages. This still allows for reasonable picture book content text for smooth reading, improving efficiency. Furthermore, even when the picture book page image is blurry, more reasonable text can be obtained for smooth reading, further enhancing efficiency. This integrates online and local electronic picture book reading. After recognizing the picture book cover and obtaining the story title, different modes are selected based on whether a locally pre-made picture book is available. Online picture book reading is more scalable, especially when the recognized image content is incomplete, utilizing the context learning ability of the image understanding model to quickly expand the story and ensure the enjoyment of reading picture books.

[0172] In some embodiments, such as Figure 6 As shown, a flowchart for obtaining the text of a picture book's inner page is provided. The locally stored collection of electronic picture books includes picture books 1, 2, and 3, etc. After first identifying the picture book names, it is determined whether there is a local electronic picture book in the locally stored collection that matches the picture book name. If it exists, the local electronic picture book is read, that is, the audio segment matching the current picture book's inner page is determined according to steps 504 and 506, or the audio segment pre-generated for the target picture book's inner page is determined, thus obtaining the audio segment matching the current picture book's inner page. If it does not exist, an online picture book is read, that is, based on the picture book's inner page image and the inner page text of at least one read picture book inner page, the inner page text of the current picture book's inner page is generated, and the inner page text is converted into audio to obtain the audio segment matching the current picture book's inner page.

[0173] In some embodiments, such as Figure 7 As shown, a system architecture diagram is provided for the picture book reading method provided in this application. The system architecture diagram includes a UI layer, a data parsing layer, and a server layer. The UI (User Interface) is the interface provided by the APP layer for displaying picture book content and narrating picture book stories. The data parsing layer implements the ability to collect and manage image data, call image understanding models, and analyze the data returned by the models. It can also implement story matching strategies and fault tolerance schemes when picture book content is not recognized. The server can provide image understanding, semantic analysis, and semantic generation services to the APP layer, and can generate audio files based on text. Among them, the image understanding model can support image analysis and semantic analysis, and can generate text based on input content. The picture book inner page text mentioned above can be generated by the image understanding model. Figure 7 The steps involved include: 1. Data acquisition: The camera captures images of the picture book cover or inner pages; 2. Image understanding: The image understanding model recognizes the acquired images and obtains the recognition results; 3. Content parsing: The recognition results are analyzed to obtain the inner page text or the picture book title; 4. Story expansion: Based on the inner page images and the inner page text of at least one read inner page, the inner page text of the current inner page is generated; 5. Semantic data; 6. Semantic generation: Speech, i.e., audio, is generated; 7. Voice feedback: Audio is output. For online reading mode, each time an audio segment is generated, the generated audio segment can be added to the picture book's audio file.

[0174] In some embodiments, this application also provides a picture book reading method, which can be executed by a reading device or a server, or jointly by the reading device and the server. The reading device can be understood as a voice-activated robot, and a picture book reading tool can be installed on the reading device. The picture book reading method provided in this application can be implemented by the reading device through the picture book reading tool, such as... Figure 9 As shown, the method includes:

[0175] Step 902: After the reading device enters the picture book reading process, acquire the picture book page image obtained by the reading device from the image capture of the currently turned picture book page.

[0176] In some embodiments, before capturing an image of the currently turned page in the picture book, the image capture device is also configured to capture an image of the picture book cover to obtain a picture book cover image; the controller is also configured to perform text recognition on the picture book cover image to obtain the picture book name.

[0177] Step 904: If there is no local electronic picture book in the collection of electronic picture books stored locally on the reading device that matches the picture book name, generate the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page, and perform audio conversion on the picture book inner page text to obtain an audio segment that matches the current picture book inner page.

[0178] Specifically, after obtaining the text of the picture book's inner pages, the reading device or server can determine whether there is a picture book inner page text in each read picture book inner page that is consistent with or substantially consistent with the text of that picture book inner page. If it is not found, the audio segment is identified; otherwise, a prompt is made to turn the page.

[0179] Step 906: Output the audio clip that matches the current page of the picture book.

[0180] In this embodiment, if no locally stored electronic picture book with a matching picture book name exists in the locally stored collection of electronic picture books, the text of the current picture book page is generated based on the picture book page image and the text of at least one read picture book page. The text is then converted into audio to obtain an audio segment that matches the current picture book page. This allows for the generation of the current picture book page text based on already read picture book pages, even when the picture book being read is not stored locally, thus ensuring a reasonable picture book content for smooth reading and improving reading efficiency.

[0181] Based on this, in some embodiments, a reading device is provided, including: an image acquisition unit configured to acquire an image of the currently turned page of a picture book after the reading device enters the picture book reading process, thereby obtaining a picture book page image; a controller configured to: generate picture book page text of the current picture book page based on the picture book page image and the picture book page text of at least one read picture book page in a locally stored collection of electronic picture books, and perform audio conversion on the picture book page text to obtain an audio segment matching the current picture book page; and an audio output unit configured to output the audio segment matching the current picture book page.

[0182] In this embodiment, if no locally stored electronic picture book with a matching picture book name exists in the locally stored collection of electronic picture books, the text of the current picture book page is generated based on the picture book page image and the text of at least one read picture book page. The text is then converted into audio to obtain an audio segment that matches the current picture book page. This allows for the generation of the current picture book page text based on already read picture book pages, even when the picture book being read is not stored locally, thus ensuring a reasonable picture book content for smooth reading and improving reading efficiency.

[0183] In some embodiments, the method of generating the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book is further configured to: generate a third prompt, the third prompt being used to instruct the generation of the picture book inner page text of the current picture book inner page using the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book; and generate the picture book inner page text of the current picture book inner page based on the third prompt, the picture book inner page image, and the picture book inner page text of at least one read picture book inner page of the picture book.

[0184] The third prompt could be something like, "You are a children's picture book explanation expert. Please identify the story plot of this page based on the picture book content of the previous chapter and the picture book image of this page," or "You are a children's picture book explanation expert. Please identify the story plot of this page based on the picture book content of the previous page and the picture book image of this page, and output the string content. If the recognition fails, return 'known'."

[0185] Specifically, the reading device or server can input the third prompt, the picture book page image, and the picture book page text of at least one read picture book page into the image understanding model to generate the picture book page text of the current picture book page.

[0186] In this embodiment, by using the third prompt and the text of at least one read page of the picture book, the content of the next page can be identified based on the previous page of the picture book's storyline. This ensures that the correct story content is obtained even when the data collection is inaccurate (blurry images, incomplete shooting angles), or the story can be expanded.

[0187] like Figure 10 The diagram illustrates a sequence of a picture book reading method. First, the user utters the voice "I want to read a picture book." The sound acquisition device captures this voice and transmits it to the controller. The controller then controls the audio output device to output the voice "Is there a picture book available?". The user then utters "Yes." The sound acquisition device captures this voice and transmits it to the controller. The controller activates the image acquisition device and controls the audio output device to output the voice "Please place the picture book in position XXX" and "Before reading the picture book, I want to look at the cover." The image acquisition device captures the picture book cover image and transmits it to the controller. The controller identifies the picture book title based on the cover image and then controls the audio output device to output the voice "Start reading the picture book." The image acquisition device captures images of the picture book's inner pages at fixed time intervals and transmits these images to the controller. The controller determines the text on the inner pages based on the images, identifies the audio segment, and controls the audio output device to output the audio segment.

[0188] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0189] Based on the same inventive concept, this application also provides a picture book reading device for implementing the picture book reading method described above. The solution provided by this device is similar to the solution described in the above method, and the specific limitations can be found in the limitations of the picture book reading method above, which will not be repeated here.

[0190] In some embodiments, a picture book reading device is provided, comprising:

[0191] The first image acquisition module is used to acquire the image of the picture book page obtained by the reading device from the image acquisition of the currently turned picture book page after the reading device enters the picture book reading process.

[0192] The first audio segment determination module is used to generate the picture book inner page text of the current picture book inner page based on the inner page image of the picture book when there is a local electronic picture book in the collection of electronic picture books stored locally on the reading device that matches the picture book name, and to determine the target picture book inner page with the highest matching degree from the local electronic picture books; when the matching degree between the picture book inner page text and the target picture book inner page does not reach the matching degree threshold, it generates the updated picture book inner page text of the current picture book inner page based on the picture book inner page text and the picture book inner page text of at least one read picture book inner page of the picture book, and performs audio conversion on the updated picture book inner page text to obtain an audio segment that matches the current picture book inner page.

[0193] The first audio output module is used to output an audio segment that matches the current page of the picture book.

[0194] In some embodiments, a picture book reading device is provided, comprising:

[0195] The second image acquisition module is used to acquire the image of the picture book page obtained by the reading device after entering the picture book reading process;

[0196] The second audio segment determination module is used to generate the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page when there is no local electronic picture book in the collection of electronic picture books stored locally on the reading device that matches the picture book name, and to perform audio conversion on the picture book inner page text to obtain an audio segment that matches the current picture book inner page.

[0197] The second audio output module is used to output an audio segment that matches the current page of the picture book.

[0198] In one embodiment, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the picture book reading method described above.

[0199] In one embodiment, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the picture book reading method described above.

[0200] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0201] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0202] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.< / storycapturebean> < / storycapturebean> < / storycapturebean>

Claims

1. A reading device, characterized in that, include: An image acquisition device is configured to acquire an image of the currently turned page of the picture book after the reading device enters the picture book reading process, thereby obtaining an image of the picture book page. The controller is configured as follows: If a local electronic picture book with a name matching the picture book exists in the locally stored collection of electronic picture books, the picture book inner page text of the current picture book inner page is generated based on the picture book inner page image, and the target picture book inner page with the highest matching degree with the picture book inner page text is determined from the local electronic picture books. If the matching degree between the picture book inner page text and the target picture book inner page does not reach the matching degree threshold, based on the picture book inner page text and the picture book inner page text of at least one read picture book inner page of the picture book, an updated picture book inner page text of the current picture book inner page is generated, and the updated picture book inner page text is converted into audio to obtain an audio segment that matches the current picture book inner page. An audio output device is configured to output an audio clip that matches the current page of the picture book.

2. The reading device according to claim 1, characterized in that, The method of generating the updated picture book inner page text of the current picture book inner page based on the picture book inner page text and at least one read picture book inner page text of the picture book is further configured as follows: Generate a first prompt, which is used to instruct the generation of the updated picture book inner page text of the current picture book inner page using the picture book inner page text and the picture book inner page text of at least one read picture book inner page of the picture book. Based on the first prompt, the picture book inner page text, and the picture book inner page text of at least one read picture book inner page, the updated picture book inner page text of the current picture book inner page is generated.

3. The reading device according to claim 1, characterized in that, The method of generating the picture book inner page text based on the picture book inner page image is further configured as follows: Generate a second prompt, which is used to instruct the recognition of text in the inner page image of the picture book; Based on the second prompt and the picture book page image, generate the picture book page text for the current picture book page.

4. The reading device according to claim 1, characterized in that, The step of determining the target picture book page with the highest matching degree to the text of the picture book page from the local electronic picture book is further configured as follows: The text on the inner pages of the picture book is split into multiple text fragments; Based on the multiple text fragments, the target picture book page with the highest matching degree with the text of the picture book page is determined from the local electronic picture book.

5. The reading device according to claim 1, characterized in that, The reading device is also configured to: If the matching degree between the text on the inner page of the picture book and the inner page of the target picture book reaches a matching degree threshold, an audio segment pre-generated for the inner page of the target picture book is determined, and an audio segment matching the current inner page of the picture book is obtained.

6. The reading device according to claim 5, characterized in that, The output audio segment that matches the current picture book page is further configured as follows: If the current reading status of the target picture book page is unread, output an audio clip that matches the current picture book page; The controller is also configured to: When outputting the audio clip, the reading status of the target picture book's inner page is updated from "unread" to "reading".

7. The reading device according to claim 6, characterized in that, The controller is also configured to: If the current reading status of the target picture book page is "read", output a page-turning prompt message; If the page-turning prompt message is displayed repeatedly, exit the picture book reading process.

8. The reading device according to any one of claims 1 to 7, characterized in that, Before capturing the image of the currently turned page in the picture book. The image acquisition device is also configured to acquire an image of the picture book cover to obtain a picture book cover image; The controller is also configured to perform text recognition on the picture book cover image to obtain the picture book title.

9. The reading device according to claim 8, characterized in that, The text recognition of the picture book cover image to obtain the picture book title is further configured as follows: The image of the picture book cover is subjected to text recognition to obtain the cover recognition text; Multiple preset texts are determined, each of which is used to characterize an inaccurate result that was not identified; The title of the picture book is determined based on the cover recognition text and the multiple preset texts.

10. The reading device according to claim 9, characterized in that, The method of determining the picture book title based on the cover recognition text and the multiple preset texts is further configured as follows: Determine the character length of the cover recognition text and determine a preset character length threshold; The picture book name is determined based on the cover recognition text, the character length, the character length threshold, and the multiple preset texts.

11. The reading device according to claim 1, characterized in that, The controller is also configured to: If no local electronic picture book with a name matching the picture book exists in the locally stored collection of electronic picture books, the picture book text of the current picture book page is generated based on the picture book page image and the picture book page text of at least one read picture book page of the picture book. The picture book page text is then converted into audio to obtain an audio segment that matches the current picture book page.

12. A reading device, characterized in that, include: An image acquisition device is configured to acquire an image of the currently turned page of the picture book after the reading device enters the picture book reading process, thereby obtaining an image of the picture book page. The controller is configured as follows: If there is no local electronic picture book in the locally stored collection that matches the picture book name, the picture book inner page text of the current picture book inner page is generated based on the inner page image of the picture book and the inner page text of at least one read inner page of the picture book, and the inner page text of the current picture book inner page is converted into audio to obtain an audio segment that matches the current inner page. An audio output device is configured to output an audio clip that matches the current page of the picture book.

13. The reading device according to claim 12, characterized in that, The method of generating the picture book inner page text of the current picture book inner page based on the picture book inner page image and the picture book inner page text of at least one read picture book inner page is further configured to: Generate a third prompt, which is used to instruct the generation of the picture book inner page text of the current picture book inner page using the picture book inner page image and the picture book inner page text of at least one read picture book inner page of the picture book. The picture book inner page text is generated based on the third prompt, the picture book inner page image, and the picture book inner page text of at least one read picture book inner page.

14. The reading device according to any one of claims 12 to 13, characterized in that, Before capturing the image of the currently turned page in the picture book. The image acquisition device is also configured to acquire an image of the picture book cover to obtain a picture book cover image; The controller is also configured to perform text recognition on the picture book cover image to obtain the picture book title.

15. A picture book reading method, characterized in that, The method includes: After the reading device enters the picture book reading process, the image of the picture book page obtained by the reading device from the image acquisition of the currently turned picture book page; If a local electronic picture book with a name matching the picture book exists in the collection of electronic picture books stored locally on the reading device, the picture book inner page text of the current picture book inner page is generated based on the picture book inner page image, and the target picture book inner page with the highest matching degree with the picture book inner page text is determined from the local electronic picture books; If the matching degree between the inner page text of the picture book and the inner page text of the target picture book does not reach the matching degree threshold, an updated inner page text of the current inner page of the picture book is generated based on the inner page text of the picture book and the inner page text of at least one read inner page of the picture book, and the updated inner page text of the picture book is converted into audio to obtain an audio segment that matches the current inner page of the picture book. Output an audio clip that matches the current page of the picture book.

16. A picture book reading method, characterized in that, The method includes: After the reading device enters the picture book reading process, the image of the picture book page obtained by the reading device from the image acquisition of the currently turned picture book page; If there is no local electronic picture book in the collection of electronic picture books stored locally on the reading device that matches the picture book name, the picture book inner page text of the current picture book inner page is generated based on the inner page image of the picture book and the inner page text of at least one read inner page of the picture book, and the inner page text of the current picture book inner page is converted into audio to obtain an audio segment that matches the current inner page of the picture book. Output an audio clip that matches the current page of the picture book.