Methods and related devices for reading picture books
By recognizing the reader's and picture book's location using a camera, the device adjusts its orientation to quickly resume reading, solving the problem of low efficiency in finding picture books after reading is interrupted and improving the user experience.
Patent Information
- Application Number
- CN202310911624.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-21
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-07-21
AI Technical Summary
During the reading of picture books, the position of the reader and the picture book becomes inconsistent due to the change in position after the device is woken up, resulting in low efficiency and poor user experience when resuming reading.
By using a camera to identify the location of the reader and the picture book, and by using voice input and camera rotation to adjust the device's orientation, the system can quickly locate the picture book near the reader and resume reading.
It improved the hit rate of picture books after reading resumed, optimized the continuity of reading and interactive experience, reduced the time of invalid location recognition, and enhanced the user's audiovisual experience.
Smart Images

Figure CN119336289B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to methods and related devices for reading picture books. Background Technology
[0002] Picture books are a type of book that primarily features illustrations with a small amount of text. Reading robots and other devices can recognize picture books and play the text aloud, or play pre-recorded audio files. This allows readers to simultaneously view the picture book and listen to the audio played by the robot, providing a dual audio-visual experience.
[0003] Besides reading picture books, robots and other devices also offer voice interaction capabilities. During picture book reading, these devices can be woken up by the user and interact with them via voice. Afterward, the robot can switch back to reading mode and continue playing the audio file. Because the wake-up time is variable, the positions of the reader and the picture book may change after returning to reading mode compared to before the robot was woken up. Providing readers with a good audiovisual experience is a direction for current and future research. Summary of the Invention
[0004] This application provides a method and related device for reading picture books, which can improve the hit rate of picture books after resuming reading, optimize the reader's reading continuity, and enhance the interactive experience of resuming reading after interruption.
[0005] The first aspect provides a method for reading picture books, which may include:
[0006] Receive a first voice input for reading a picture book; in response to the first voice input, acquire a first image through a first camera, the first image containing the face of the first user who made the first voice input; acquire a second image through a second camera in a first direction, the second image including the first picture book, and play the audio of the first picture book;
[0007] Receive a second voice input containing a wake word; in response to the second voice input, stop playing the audio of the first picture book, rotate to face the second user who issued the second voice input, and engage in voice interaction with the second user; acquire a third image through the first camera in the second direction, the third image being an image of the first user's face acquired through the first camera during the voice interaction with the second user;
[0008] Receive a third voice input for reading a picture book; in response to the third voice input, rotate to a second direction; acquire a fourth image via a first camera in the second direction; if the fourth image contains the face of the first user and the angle between the second direction and the first direction is less than a threshold, rotate to the first direction; acquire a fifth image via a second camera in the first direction; if the fifth image contains the first picture book, continue playing the audio of the first picture book.
[0009] The second aspect provides a method for reading picture books, which may include:
[0010] Receive a first voice input for reading a picture book; in response to the first voice input, acquire a first image through a first camera, the first image containing the face of the first user who made the first voice input, and face a second direction when acquiring the first image; acquire a second image through a second camera in the first direction, the second image including the first picture book, and play the audio of the first picture book;
[0011] Receive a second voice input containing a wake word; in response to the second voice input, stop playing the audio of the first picture book, rotate to face the second user who issued the second voice input, and engage in voice interaction with the second user. During the voice interaction with the second user, the image acquired by the first camera does not contain the face of the first user.
[0012] Receive a third voice input for reading a picture book; in response to the third voice input, rotate to a second direction; acquire a fourth image via a first camera in the second direction; if the fourth image contains the face of the first user and the angle between the second direction and the first direction is less than a threshold, rotate to the first direction; acquire a fifth image via a second camera in the first direction; if the fifth image contains the first picture book, continue playing the audio of the first picture book.
[0013] The first approach applies to scenarios where the device has seen the first user after being woken up.
[0014] The second approach applies to scenarios where the device has not seen the first user after being woken up.
[0015] By employing either the first or second method, when resuming reading a picture book after an interruption, the system first locates the reader's historical position and then uses the reader as the primary clue to find the picture book. This ensures that reading continues only when the reader and the picture book are close together, and also improves the speed at which the device finds the book. Essentially, this not only increases the success rate of finding the picture book after resuming reading but also optimizes the reader's reading continuity and enhances the interactive experience of resuming reading after an interruption.
[0016] In conjunction with the first or second aspect, in some embodiments, the second image may be any image containing the first picture book acquired by the second camera before receiving the second voice input, or it may be the most recently acquired image containing the first picture book acquired by the second camera before receiving the second voice input.
[0017] In conjunction with the first aspect, in some embodiments, the third image may be any image containing the face of the first user acquired by the first camera during the voice interaction with the second user, or it may be the most recently acquired image containing the face of the first user acquired by the first camera during the voice interaction with the second user.
[0018] In conjunction with the first or second aspect, in some implementations, if the fifth image does not contain the first picture book, a first prompt message is output, which prompts the user to place the first picture book.
[0019] In conjunction with the first or second aspect, in some implementations, if the fourth image contains the face of the first user and the angle between the second direction and the first direction is greater than a threshold, a sixth image is acquired in the second direction via a second camera; if the sixth image contains the first picture book, the audio of the first picture book continues to be played.
[0020] In conjunction with the previous implementation method, if the sixth image does not contain the first picture book, a second prompt message is output, which prompts the user to place the first picture book.
[0021] In conjunction with the first or second aspect, in some embodiments, where the fourth image does not contain the face of the first user,
[0022] If a seventh image is acquired via the first camera in a third-party direction, and the seventh image contains the face of the first user, then an eighth image is acquired via the second camera in a third-party direction; if the eighth image contains the first picture book, the audio of the first picture book continues to play.
[0023] In conjunction with the previous implementation method, if the eighth image does not contain the first picture book, a third prompt message is output, which is used to prompt the user to place the first picture book.
[0024] In conjunction with the first or second aspect, in some implementations, if the first camera does not acquire an image containing the first user's face within a first time period when the fourth image does not contain the first user's face, the device enters a sleep state or a power-off state.
[0025] In conjunction with the first or second aspect, in some embodiments, before acquiring a second image via the second camera in the first direction, the method further includes: turning off the first camera.
[0026] In conjunction with the first or second aspect, in some implementations, the second camera is a wide-angle camera.
[0027] In conjunction with the first or second aspect, in some implementations, the servo motor can be rotated to a second direction by adjusting its rotation angle.
[0028] A third aspect provides an apparatus comprising: a memory, one or more processors; the memory being coupled to the one or more processors, the memory being used to store program code, the program code including instructions, the one or more processors invoking the instructions to cause the apparatus to perform a method as described in the first aspect, or the second aspect, or any embodiment of the first aspect, or any embodiment of the second aspect.
[0029] A fourth aspect provides a readable storage medium including instructions that, when executed on a device, cause the device to perform a method as described in the first aspect, or the second aspect, or any embodiment of the first aspect, or any embodiment of the second aspect.
[0030] The fifth aspect provides a program product that, when run on a device, causes the device to perform a method as described in the first aspect, or the second aspect, or any implementation of the first aspect, or any implementation of the second aspect.
[0031] A sixth aspect provides a chip system including at least one processor for implementing a method as described in the first aspect, or the second aspect, or any implementation of the first aspect, or any implementation of the second aspect. Attached Figure Description
[0032] Figure 1 A device diagram in robot form provided in the embodiments of this application;
[0033] Figure 2 The scenario in which the process of reading a picture book on the device provided in the embodiments of this application is interrupted;
[0034] Figure 3 A flowchart illustrating a method for reading picture books provided in an embodiment of this application;
[0035] Figure 4 This application provides schematic diagrams illustrating several scenarios in which the device is located, as shown in the embodiments of this application.
[0036] Figure 5 This is a hardware structure block diagram of the device provided in the embodiments of this application;
[0037] Figure 6 A software structure block diagram of the device provided in the embodiments of this application. Detailed Implementation
[0038] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0039] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0040] The term "user interface (UI)" used in the following embodiments of this application refers to the medium interface through which an application or operating system interacts and exchanges information with the user. It realizes the conversion between the internal form of information and the form that the user can accept. The user interface is source code written in a specific computer language such as Java or Extensible Markup Language (XML). The interface source code is parsed and rendered on the device, ultimately presenting content that the user can recognize. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be visible interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets displayed on the device's screen.
[0041] The method for reading picture books provided in this application is applied to a device. This device can take various product forms, such as a robot, smart speaker, smart lamp, smart furniture, smart transportation, etc. The device is equipped with a camera that can rotate, and the camera's position in space changes after rotation.
[0042] Figure 1 An example of a device 100 in the form of a robot is shown.
[0043] like Figure 1As shown, the robot's head may include cameras, such as camera 193-1 and camera 193-2. Both camera 193-1 and camera 193-2 can be used to acquire images within their own field of view. Camera 193-1 can be a wide-angle camera, used to acquire images containing picture books; camera 193-2 can be used to acquire images containing human faces. The default initial orientation of camera 193-1 can be below the horizontal line and at a certain angle (e.g., 20 degrees) to the horizontal line, which facilitates camera 193-1 acquiring images of picture books laid flat on a table. The default initial orientation of camera 193-2 can be on the horizontal line. The robot's head can rotate to the left, right, up, down, or other directions using servos. After the robot's head position changes, the positions of cameras 193-1 and 193-2 in space change, and the images they can acquire also change accordingly. In other words, the robot can adjust the field of view of cameras 193-1 and 193-2 using servos.
[0044] Of course, in some other implementations, the robot may include only one camera.
[0045] The robot may also include a microphone array for receiving sound so that the robot can determine the location of the sound source.
[0046] Figure 2 An example is shown where the process of reading a picture book on a device is interrupted.
[0047] like Figure 2 As shown, the reader ( Figure 2 Both user1 and the picture book are located at position P1.
[0048] First, device 100 faces P1 and uses its head-mounted camera 193-1 to acquire an image containing the picture book and play audio (speech synthesized from the text in the picture book, or a pre-recorded audio file of the picture book). When device 100 detects the presence of a picture book, it determines the current location of the picture book, P1, and can continuously refresh the location of the picture book to determine the last location of the picture book it has detected.
[0049] Then, while the device 100 is playing audio, the waker located at position P2 ( Figure 2 The user2 in the context of the device outputs a wake-up word (e.g., "Xiaoyi Xiaoyi"). The device 100 can receive voice input containing the wake-up word, determine the location P2 of the person issuing the voice input, turn its head to face P2, and simultaneously activate the voice assistant in response to the wake-up word. Voice interaction is then conducted between the voice assistant and the person who activated the wake-up word. After being activated, the device 100 can turn off camera 193-1 and activate camera 193-2 to locate the person who activated the wake-up word, allowing it to face that person during the conversation.
[0050] Subsequently, if the waker outputs voice input instructing device 100 to read the picture book (such as "read the picture book") or does not output any command for an extended period, device 100 resumes reading mode. Specifically, reading mode is resumed as follows: device 100 turns off camera 193-2 and turns on camera 193-1; it first faces P2 and acquires an image through camera 193-1, and identifies whether a picture book exists based on the image; if it exists, it plays the corresponding audio; if it does not exist, it turns to the last position of the picture book, P1, and identifies whether a picture book exists again.
[0051] Figure 2 The picture book reading process shown involves the following concepts:
[0052] 1. Reader
[0053] A reader refers to a user who needs to read picture books, such as a child. There can be one or more readers.
[0054] 2. The Awakener
[0055] The awakener refers to the user who sends the wake word. The awakener and the reader can be the same user or different users.
[0056] 3. Voice Assistant
[0057] A voice assistant is a feature provided by a device that allows the device to receive voice input from the user, use speech recognition technology and natural language processing technology (such as semantic understanding) to identify the user's intent, and respond accordingly, such as providing voice answers, launching applications, changing device settings, etc. The device can activate the voice assistant after receiving a wake word (which can be a wake word bound to user input).
[0058] Figure 2 The scenario shown, where reading a picture book is interrupted and then resumed, has the following shortcomings: If the picture book's position hasn't moved, device 100 first faces P2 and cannot recognize the book at P2, then it turns to P1 to recognize it. If one recognition cycle takes about 1.5 seconds, resuming reading would take 4 seconds or more. Since the picture book's position hasn't moved, the process of device 100 facing P2 and recognizing the book is meaningless to the user, and this time is wasted. Furthermore, if the reader has already left P1, device 100 turning to P1 to continue reading the picture book provides a poor user experience.
[0059] The picture book reading method provided in the following embodiments of this application, when resuming reading after an interruption, first locates the original location of the reader and then uses the reader as the primary clue to find the picture book. This ensures that reading can continue only when the reader and the picture book are close together, and also improves the speed at which the device finds the picture book. In essence, this not only increases the picture book accuracy rate after resuming reading but also optimizes the reader's reading continuity and improves the interactive experience of resuming reading after an interruption.
[0060] refer to Figure 3 , Figure 3 The flowchart of the method for reading picture books provided in this application is illustrated by example. (Reference) Figure 3 The method may include the following steps:
[0061] Phase 1: Reading picture books
[0062] S11, device 100 activates camera 193-2 and acquires images through camera 193-2.
[0063] Optionally, camera 193-2 can be kept on after device 100 is powered on.
[0064] S12, device 100 receives voice input for reading picture books.
[0065] Device 100 can receive voice input via a microphone or other sound pickup device. The voice input used for reading picture books can be preset by device 100 or set by the user, for example, the voice input can be "read picture book".
[0066] After receiving voice input for reading picture books, device 100 can enter reading mode. Reading mode can refer to device 100 executing S13-S15.
[0067] S13, the device 100 acquires the reader's facial image through the camera 193-2 and determines the reader's location P2.
[0068] The reader refers to the user in S12 who provides voice input for reading picture books.
[0069] After receiving voice input for reading picture books, device 100 can respond to the voice input and locate the reader.
[0070] Specifically, device 100 can employ microphone array sound source localization technology to determine the reader's direction based on the voice input given by the reader in S12, then rotate to that direction, acquire an image in that direction via camera 193-2, and identify the face in the image as the reader's face. Device 100 can rotate its direction by adjusting its posture, head direction, or servo motor rotation angle.
[0071] If the image captured by camera 193-2 in the direction of the reader contains multiple faces, device 100 can first extract the voiceprint from the voice input, then find the reader's face from the pre-stored correspondence between the voiceprint and the user's face, and then find the reader's face from the multiple faces in the image.
[0072] The reader's position reflects the reader's relative position on the device 100, and can be represented by one or more of the following data: the pose of the device 100 when the face image is acquired (such as the pose of a robot), the direction that the device 100 (such as camera 193-2) is facing, the state of the servo motor, etc.
[0073] S14, device 100 turns off camera 193-2, turns on camera 193-1, and acquires images through camera 193-1.
[0074] The camera 193-1 can be a wide-angle camera, which is convenient for capturing a wide field of view and is used for subsequent identification of picture books.
[0075] S15, device 100 recognizes the picture book in the image captured by camera 193-1, and interacts with the reader based on the recognized picture book.
[0076] After activating camera 193-1, device 100 can rotate to change the field of view of camera 193-1 until camera 193-1 captures an image containing the picture book. When the picture book is within the field of view of camera 193-1, device 100 can acquire an image containing the picture book through camera 193-1. If device 100 does not acquire an image containing the picture book, it can output a prompt message to remind the user to place the picture book within the field of view of camera 193-1 until camera 193-1 acquires an image containing the picture book. This prompt message can be either voice or text displayed on the screen.
[0077] Device 100 can identify picture books in an image using picture book recognition technology. This application embodiment does not limit the scope of this picture book recognition technology. For example, a QR code may be printed on the picture book, and device 100 can identify the QR code in the acquired image to obtain picture book information (such as identification).
[0078] The device 100 can use the following methods to interact with the reader based on the recognized picture book:
[0079] 1. Device 100 plays the audio corresponding to the picture book. For example, device 100 finds the audio file corresponding to the picture book from multiple pre-made audio files of picture books stored in the memory and plays it. Alternatively, device 100 can recognize the text in the acquired image using optical character recognition (OCR), and then synthesize these texts into speech and play it.
[0080] 2. Device 100 plays interactive content corresponding to the picture book. After receiving feedback from the reader, it plays subsequent interactive content based on the reader's feedback. For example, the robot-shaped device 100 can play questions corresponding to the picture book. The reader can answer by touching the robot's head, left hand, right hand, belly, etc. Sensors at these locations on the robot can detect the reader's touch operation and continue playing further questions based on the reader's answer.
[0081] 3. Device 100 can engage in real-time voice interaction with the reader based on the picture book. This voice interaction process is based on the content of the picture book. Device 100 can provide different responses based on different voice inputs from the user.
[0082] During the interaction between the device 100 and the reader based on the recognized picture book, if the position of the picture book changes, the device 100 can turn to face the new position to locate the picture book using the camera 193-1. In other words, the position of the picture book may be updated in real time during this process.
[0083] Phase Two, Interruption of Reading
[0084] S16, Device 100 receives voice input containing a wake word from the waker.
[0085] Device 100 can receive voice input via a microphone or other sound pickup device. The wake word can be preset by device 100 or set by the user.
[0086] Device 100 can respond to a wake word in voice input, activate the voice assistant, and stop interaction with the reader based on the recognized picture book, such as stopping the playback of the picture book's audio.
[0087] S17, Device 100 determines the location of the picture book P1.
[0088] In some implementations, the device 100 can periodically or non-periodically determine the location of the picture book during the picture book reading process in Phase 1, and record the last determined location of the picture book in Phase 1 as P1 after receiving a wake word. This means that S17 can be executed at the end of Phase 1 and before Phase 2. In other implementations, any picture book location determined during the picture book reading process in Phase 1 can be recorded as P1. S17 can then be executed before S16.
[0089] In other implementations, the device 100 may determine the current location of the picture book and record it as P1 after receiving the wake word.
[0090] The location of the picture book reflects its relative position on the device 100, which can be reflected by one or more of the following data: the pose of the device 100 (such as the pose of a robot) when reading the picture book, the direction that the device 100 (such as camera 193-1) is facing, the status of the servo motor, etc.
[0091] S18, Device 100 turns off camera 193-1 and turns on camera 193-2.
[0092] S19, Device 100 rotates to face the waker and engages in voice interaction with the waker.
[0093] Device 100 can employ microphone array sound source localization technology to determine the direction of the waker based on the voice input from the waker, and can rotate to that direction. Voice interaction refers to the process where the waker outputs voice input, and device 100 responds to that voice input through a voice assistant. During voice interaction, the waker can move, and device 100 can continuously update the waker's location and adjust its direction in real time to locate the waker. Device 100 facing the waker makes the voice interaction experience more user-friendly.
[0094] S20, during voice interaction, if the image acquired by camera 193-2 contains the reader's face image, then the reader's location is determined, and the reader's location P2 is refreshed to the reader's location when camera 193-2 acquired the reader's face image.
[0095] Device 100 can use facial recognition technology to determine whether the image acquired by camera 193-2 contains the reader's face in S18.
[0096] During the voice interaction in S19, the device 100 can rotate to locate the person waking up, thus enabling the camera 193-2 to capture images from different directions. If the reader moves during the voice interaction and the new location is still within the field of view of the device 100, the device 100 can capture and update the location.
[0097] In some other implementations, the reader's position P2 can be refreshed to any position of the reader's face image acquired by the camera 193-2 during the voice interaction process.
[0098] Of course, during voice interaction, there may be situations where the image acquired by camera 193-2 does not contain the reader's face. In this case, S20 does not need to be executed. If, during voice interaction, the image acquired by camera 193-2 does not contain the reader's face, that is, device 100 does not capture the reader's face, then P2 will still be the original data determined in S16.
[0099] Phase Three: Resumption of Reading
[0100] S21, device 100 receives voice input for reading picture books, or, for a period of time, does not receive any voice input.
[0101] It should be understood that the voice input used for reading picture books in S21 and the voice input used for reading picture books in S12 can be the same. For example, the voice input in S21 can be "Read the picture book", "Let's read the picture book", "Continue reading the picture book", etc.
[0102] The voice input used for reading picture books in S21 can be issued by the waker, the reader, or other users; there are no restrictions on this.
[0103] S22, device 100 rotates to the reader-facing position P2.
[0104] The device can face the reader's position P2 by adjusting its posture, head direction, and servo motor rotation angle. This reader's position P2 can be the reader's location determined by device 100 in S13, or any reader's location determined during the voice interaction process in S20, or the latest or last determined reader's location during the voice interaction process in S20.
[0105] S23, when the device 100 is facing the reader position P2, it determines whether the image acquired by the camera 193-2 contains the reader's face.
[0106] If yes, then execute S24; otherwise, execute S30.
[0107] S24, Device 100 determines whether the angle θ between the positions of P2 and P1 is less than the threshold.
[0108] The angle between P1 and P2 refers to the angle between the direction in which device 100 faces P1 and the direction in which it faces P2.
[0109] The threshold can be determined based on the angle θ within the field of view of camera 193-1, and can be less than or equal to that angle θ. This threshold can also be continuously updated by device 100 based on actual conditions. For example, the threshold can be 20°.
[0110] If yes, then execute S29; otherwise, execute S25.
[0111] If the angle θ between the positions of P2 and P1 is less than the threshold, assuming the picture book is still at the position P1 determined in S17, it means that the picture book and the reader are not far apart. Then, executing S29, facing the picture book, can more accurately identify the picture book.
[0112] If the angle θ between the positions of P2 and P1 is greater than the threshold, assuming that the picture book is not at the position P1 determined in S17, but moves with the reader, then by executing S25-S26 and facing the reader, the picture book that moves with the reader can be identified.
[0113] S25, Device 100 turns off camera 193-2 and turns on camera 193-1.
[0114] S26, Device 100 determines whether a picture book has been identified from the image acquired by camera 193-1.
[0115] Device 100 can be set to a preset duration (e.g., 2 seconds) and determine whether the picture book is recognized within the preset duration after the camera 193-1 is activated. This preset duration can be referred to as the first duration.
[0116] If yes, then execute S27; otherwise, execute S28.
[0117] S27, Device 100 is based on the recognized picture book and reader interaction.
[0118] For information on identifying picture books and how to interact with readers based on those books, please refer to S15.
[0119] In some implementations, in S26, device 100 can further determine whether the same picture book from stage one (S15) has been recognized. If so, the reading progress at the moment of interruption in stage one resumes audio playback. Correspondingly, in stage two (S16), when device 100 receives the wake word, it can record the reading progress so that the reading progress can be resumed later. This scheme allows the reader to continue reading the same picture book continuously and smoothly. Of course, if device 100 does not recognize the same picture book from stage one (S15) in S26, it can prompt the user to place the same picture book, or it can directly read the newly recognized picture book.
[0120] In other implementations, if device 100 recognizes a picture book in S26, but the picture book is not the same as the one in phase one S15, device 100 may also interact with the reader based on the newly recognized picture book.
[0121] S28, Device 100 outputs a prompt message to remind the user to place the picture book.
[0122] The prompt message can be displayed as voice output or text on a screen.
[0123] This prompt message can be used to remind the user to place the picture book near the device 100, for example, within the current field of view of the device 100. For example, the prompt message can be implemented as a voice output saying "Please place the picture book directly in front of me".
[0124] In some implementations, the prompt message may further suggest a specific picture book to the user, such as the same picture book used for the user placement in stage one S15.
[0125] If the user places a picture book based on the prompt, the device 100 can recognize the picture book and continue reading.
[0126] S29, the device 100 rotates to face the position P1 where the picture book is located.
[0127] The picture book location P1 here can be any picture book location determined by device 100 during the picture book reading process in stage one, or it can be the latest or last picture book location determined by device 100 in stage one.
[0128] After executing S29, the device returns to S25.
[0129] If device 100 enters phase three and executes steps S21-S24, S29, and S25-S27 sequentially, optionally, a prompt message can be output before step S27 to prompt the reader to approach the picture book. This prompt message can be either voice output or text displayed on the screen. When the reader approaches the picture book, they can listen to the audio played by device 100 while simultaneously viewing the picture book, thus receiving information from both visual and auditory senses.
[0130] S30, Device 100 is looking for readers.
[0131] The device 100 can adjust its facing direction by changing its posture, head direction, and servo motor rotation angle. This allows it to acquire different images via camera 193-2 and determine whether these images contain the reader's face. If they do, the device identifies the reader.
[0132] There are various ways and strategies for adjusting the orientation of device 100, which are not limited here. For example, device 100 can rotate its head 360° horizontally from facing P2 to find the reader; or it can first adjust to the position where it faces the reader for the second to last time, then adjust to the position where it faces the reader for the third to last time, and so on.
[0133] S31, Device 100 determines whether a reader has been found.
[0134] After starting the search, if the device 100 finds a reader, it can immediately execute S31 and determine the result as yes.
[0135] If no reader is found after a period of time after the device 100 starts searching, it can immediately execute S31 and determine the result as negative.
[0136] If the judgment result of S31 is yes, then device 100 returns to S25, or device 100 updates P2 to the position of the reader when the reader is found and returns to S24.
[0137] If the result of S31 is negative, then S32 is executed.
[0138] S32, Device 100 exits reading mode.
[0139] Exiting reading mode may include device 100 turning off camera 193-2.
[0140] In some implementations, exiting reading mode may include the device 100 entering a sleep state or a power-off state.
[0141] The technical effects of the picture book reading method provided in this application will be explained below, using various scenarios of actual picture book reading as examples.
[0142] refer to Figure 4 , Figure 4 Five scenarios in which a robot-shaped device exists are illustrated.
[0143] The main difference between the various scenarios lies in the positions of the reader and the picture book during the third stage of resuming reading. (Not limited to...) Figure 4 These are the five scenarios mentioned above. In practice, many more scenarios may be included, which will not be listed here.
[0144] exist Figure 4 In the scenarios shown, user1 is the reader and user2 is the arouser. Readers in stage one are represented by dashed lines, and readers in stage three are represented by solid lines.
[0145] Scene 1
[0146] Phase 1: User1 and the picture book are located at L1, and the robot faces L1.
[0147] Phase Two: User2 wakes up the robot at L2, and the robot faces L2.
[0148] In Phase 3, user1 and the picture book are still located in L1.
[0149] In scenario 1, the robot determines the location of the picture book as L1 (P1), and the robot determines the reader's final location as L1 (P2) in phase one. The robot does not update the reader's location during the wake-up process.
[0150] Implementing the method of this application in Scenario 1, after entering Phase 3, the robot sequentially executes S21-S24, S29, and S25-S27. Therefore, facing P2 (i.e., L1), the robot can both locate the reader and recognize the picture book, interacting with the reader based on the picture book. It is evident that in Scenario 1, the robot can quickly locate the picture book and resume reading, optimizing the user's reading continuity. Furthermore, having the picture book and the reader in the same location after resuming reading ensures the reader's audiovisual experience.
[0151] Scene 2
[0152] Phase 1: User1 and the picture book are located at L1, and the robot faces L1.
[0153] Phase Two: User2 wakes up the robot at L2, and the robot faces L2.
[0154] In Phase 3, the picture book remains in L1, user1 moves to L3, and the positional angle θ between the robot and L1 and L3 is less than the threshold.
[0155] In scenario 2, the robot determines the location of the picture book as L1 (P1). During the wake-up process, the robot sees the reader and the latest determined location of the reader is L3 (P2).
[0156] Implementing the method of this application in Scenario 2, after entering Phase 3, the robot sequentially executes S21-S24, S29, and S25-S27. The robot initially faces P2 (i.e., L3) and can locate the reader at P2. However, since the angle θ between P2 (i.e., L3) and P1 (i.e., L1) is less than a threshold, it turns to P1 (i.e., L1), and then can accurately identify the picture book and interact with the reader based on the picture book. It is evident that in Scenario 2, after resuming reading, the picture book and the reader are not far apart, allowing the robot to quickly find the picture book and resume reading, optimizing the user's reading continuity. Furthermore, the short distance between the picture book and the reader after resuming reading also ensures the reader's audiovisual experience.
[0157] Scene 3
[0158] Phase 1: User1 and the picture book are located at L1, and the robot faces L1.
[0159] Phase Two: User2 wakes up the robot at L2, and the robot faces L2.
[0160] In Phase 3, both the picture book and user1 move to L4, and the positional angle θ between the robot and L1 and L4 is greater than the threshold.
[0161] In scenario 3, the robot determines the location of the picture book as L1 (P1). During the wake-up process, the robot sees the reader and the latest determined location of the reader is L4 (P2).
[0162] Implementing the method of this application in scenario 3, after entering stage three, the robot executes S21-S27 sequentially. The robot first faces P2 (i.e., L4) and can find the reader at P2. Since the angle θ between P2 (i.e., L4) and P1 (i.e., L1) is greater than the threshold, the picture book can be directly identified at P2, and the robot can interact with the reader based on the picture book. It can be seen that in scenario 3, even when the reader moves with the picture book and moves at a large angle, the robot can quickly find the picture book and resume reading, optimizing the user's reading continuity.
[0163] Scene 4
[0164] Phase 1: User1 and the picture book are located at L1, and the robot faces L1.
[0165] Phase Two: User2 wakes up the robot at L2, and the robot faces L2.
[0166] In Phase 3, the picture book remains at L1, while user1 moves to L5.
[0167] In scenario 4, the robot determines the location of the picture book as L1 (P1). The robot does not see the reader during the wake-up process, and the latest determined location of the reader, P2, is still L1.
[0168] Implementing the method of this application in Scenario 4, after entering Phase 3, the robot sequentially executes S21-S23, S30, S31, S25-S27, or sequentially executes S21-S23, S30, S31, S25, S26, S28. The robot first faces P2 (i.e., L1) and, if it does not find the reader at P2, it turns to search for the reader. After finding the reader at L5, it then determines whether a picture book can be recognized at that location. If the picture book can be recognized, it interacts with the reader based on the picture book; otherwise, it prompts the user to place the picture book. Therefore, in Scenario 4, the robot can quickly find the reader, thereby resuming reading and optimizing the user's reading continuity.
[0169] Scene 5
[0170] Phase 1: User1 and the picture book are located at L1, and the robot faces L1.
[0171] Phase Two: User2 wakes up the robot at L2, and the robot faces L2.
[0172] In Phase 3, the picture book remains at L1, while user1 moves to L6.
[0173] In scenario 5, the robot determines the location of the picture book as L1 (P1). The robot does not see the reader during the wake-up process, and the latest determined location of the reader, P2, is still L1.
[0174] Implementing the method of this application in scenario 5, after entering stage three, the robot sequentially executes S21-S23, S30, S31, and S32. The robot first faces P2 (i.e., L1) and, finding no reader at P2, turns to search for the reader. However, since the reader's location, L6, is outside the robot's field of view, the robot cannot find the reader. Therefore, in scenario 5, the robot does not see the reader during the wake-up process and fails to find the reader after resuming reading, indicating that the reader has moved away. The robot can then exit the reading process, avoiding a rigid continuation of reading the picture book and providing the user with a more flexible experience.
[0175] based on Figure 3 In the method shown in this embodiment, camera 193-2 can be referred to as the first camera, and camera 193-1 can be referred to as the second camera. The reader can be referred to as the first user, and the waker can be referred to as the second user. The picture book recognized by device 100 can be referred to as the first picture book.
[0176] In Phase One:
[0177] The voice input used for reading picture books in S12 can be referred to as the first voice input. In response to the first voice input, the image containing the reader's face acquired by the camera 193-2 in S13 can be referred to as the first image.
[0178] The image containing the picture book acquired by camera 193-1 in S15 can be referred to as the second image. The second image can be any image containing the picture book acquired by camera 193-1 in S15, or it can be the last or latest image containing the picture book acquired in phase one. The direction in which device 100 faces when acquiring the second image can be referred to as the first direction.
[0179] In Phase Two:
[0180] S16 voice input containing a wake word can be referred to as second voice input.
[0181] If the image acquired by camera 193-2 does not contain the reader's face during voice interaction, i.e. S20 is not executed, then the direction in which device 100 faces when acquiring the first image can be referred to as the second direction.
[0182] If, during the voice interaction, the image acquired by camera 193-2 contains the reader's face (i.e., S20 is executed), then the image containing the reader's face acquired by device 100 through camera 193-2 during this voice interaction can be referred to as the third image. The third image can be any image containing the reader's face acquired by camera 193-2 during this voice interaction, or it can be the latest or last image containing the reader's face acquired by camera 193-2. The direction in which device 100 is facing when acquiring the third image can be referred to as the second direction.
[0183] In Phase Three:
[0184] The voice input used for reading picture books in S21 can be referred to as third voice input.
[0185] The image acquired by camera 193-2 in S23 can be referred to as the fourth image.
[0186] If S21-S24, S29, S25-S26, and S27 / S28 are executed sequentially after entering stage three, then the image acquired by camera 193-1 in S25 can be referred to as the fifth image. If the result of S26 is negative, then the prompt information input in S28 can be referred to as the first prompt information.
[0187] If S21-S26 and S27 / S28 are executed sequentially after entering stage three, then the image acquired by camera 193-1 in S25 can be referred to as the sixth image. If the result of S26 is negative, then the prompt information input in S28 can be referred to as the second prompt information.
[0188] If, after entering stage three, steps S21-S23, S30, S31, S25-S26, and S27 / S28 are executed sequentially, then the image containing the reader's face acquired by camera 193-2 in S30 can be called the seventh image, and the direction the reader is facing when acquiring the seventh image can be called the third direction. Subsequently, the image acquired by camera 193-1 in S25 in this third direction can be called the eighth image. If the result of S26 is negative, then the prompt information input in S28 can be called the third prompt information.
[0189] based on Figure 3 In addition to the methods for reading picture books, this application also provides the following extended solutions:
[0190] Extended solutions
[0191] Device 100 can acquire images through a camera, using these images to both identify picture books and to identify or locate readers. Essentially, cameras 193-1 and 193-2 are combined into a single camera.
[0192] refer to Figure 5 , Figure 5 This is a hardware structure diagram of the device 100 provided in the embodiments of this application.
[0193] like Figure 5 As shown, device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a sensor module 180, buttons 190, an indicator 192, a camera 193, a display screen 194, etc. The sensor module 180 may include a gyroscope sensor 180B, an accelerometer sensor 180E, etc.
[0194] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on device 100. In other embodiments of this application, device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0195] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0196] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0197] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0198] The charging management module 140 receives charging input from the charger. The power management module 141 connects to the battery 142, and the charging management module 140 connects to the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, display 194, camera 193, and wireless communication module 160, etc.
[0199] The wireless communication function of device 100 can be implemented through an antenna, wireless communication module 160, modem processor, and baseband processor.
[0200] Device 100 implements display functions through GPU, display screen 194, and application processor.
[0201] Device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0202] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0203] Camera 193 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP (Internet Service Provider) for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP (Digital Signal Processor) for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0204] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0205] Device 100 can implement audio functions, such as music playback and recording, through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor.
[0206] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0207] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Device 100 can listen to music or make hands-free calls through the speaker 170A.
[0208] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Device 100 may have at least one microphone 170C. In some embodiments, device 100 may have two microphones 170C, which, in addition to receiving sound signals, can also perform noise reduction. In other embodiments, device 100 may have three, four, or more microphones 170C, which can receive sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0209] The gyroscope sensor 180B can be used to determine the motion attitude of the device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the device 100's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the device 100 through reverse movement, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.
[0210] The 180E accelerometer sensor can detect the magnitude of acceleration of device 100 in various directions (typically three axes). When device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify device posture and is applicable to screen orientation switching, pedometers, and other applications.
[0211] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Device 100 can receive button input and generate key signal inputs related to user settings and function control of device 100.
[0212] In the embodiments of this application:
[0213] The number of cameras 193 can be one or more. For example, device 100 may include camera 193-1 and camera 193-2, wherein camera 193-1 can be a wide-angle camera. The image acquired by camera 193-1 can be used by processor 110 to identify and locate picture books, and the image acquired by camera 193-2 can be used by processor 110 to identify and locate readers.
[0214] Device 100 may also include a servo motor ( Figure 5 (Not shown in the image), the servo can adjust the rotation angle to adjust the direction in which the camera 193 is facing. In other embodiments, the servo can also be implemented as other types of steering devices for adjusting the direction in which the camera 193 is facing.
[0215] Microphone 170C can be used to receive voice input from the waker, such as voice input containing a wake word, voice input during voice interaction, etc. Microphone 170C can be implemented as a microphone array, and the sound received by the microphone array is used by processor 110 to determine the direction of the waker using the principle of sound source localization.
[0216] The speaker 170A can be used to play audio corresponding to picture books, and can also be used to output voice prompts.
[0217] The internal memory 121 can be used to store preset picture book data, such as picture book information (e.g., identifier) and the corresponding audio file of the picture book. In some embodiments, the wireless communication module 160, etc., can also be used to obtain the audio file corresponding to the picture book from the network.
[0218] The internal memory 121 is also used to store the implementation code of the method for reading picture books provided in this application. The processor 110 is used to read the code and implement the processing logic of the method. For example, the processor 110 can be used to identify the picture book in the image in stage one, identify whether the voice input contains a wake word in stage two, determine the position P1 of the picture book, identify the reader's face image, determine the direction in which the device 100 is facing, etc.
[0219] For details on the specific operations performed by each component of device 100, please refer to [the relevant documentation / reference]. Figure 3 The relevant descriptions will not be elaborated here.
[0220] The software system of Device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture.
[0221] Figure 6 This is a software structure block diagram of the device 100 according to an embodiment of this application.
[0222] refer to Figure 6 The device 100 may include: a face detection module, a picture book recognition module, a picture book position determination module, a picture book tracking module, a face tracking module, and a direction control module.
[0223] The face detection module is used to acquire the image stream reported by the camera (such as camera 193-2) and identify whether the image stream contains a face.
[0224] The picture book recognition module is used to acquire the image stream reported by the camera (such as camera 193-1), extract the picture book features in the image stream, and recognize the picture book information.
[0225] The picture book location determination module is used to determine the location P1 of the picture book after the picture book recognition module has identified the picture book.
[0226] The picture book tracking module is used to determine the direction in which the device 100 is facing, and whether to search for the reader, based on the detection results of the face detection module and the recognition results of the picture book recognition module.
[0227] The face tracking module is used to locate the reader based on the decision results of the picture book tracking module. This includes deciding which direction the device 100 should look for the reader, determining whether the image reported by the camera contains the reader's face, and updating the reader's location P2.
[0228] The orientation control module is used to adjust the rotation angle based on the decision results of the picture book tracking module or the face tracking module, thereby adjusting the orientation of the camera 193.
[0229] For details on the specific functions of each of the above modules, please refer to [link / reference]. Figure 3 The steps of the method shown will not be repeated here.
[0230] It should be understood that each step in the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.
[0231] This application also provides an apparatus that may include a memory and a processor. The memory may be used to store a program; the processor may be used to invoke the program in the memory, causing the apparatus to execute the method executed on the device side in any of the above embodiments.
[0232] This application also provides a chip system including at least one processor for implementing the functions involved on the device side in any of the above embodiments.
[0233] In one possible design, the chip system also includes a memory for storing program instructions and data, which may be located within or outside the processor.
[0234] The chip system can consist of chips or include chips and other discrete components.
[0235] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.
[0236] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application embodiment does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application embodiment does not specifically limit the type of memory or the arrangement of the memory and processor.
[0237] For example, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0238] This application also provides a program product comprising: a program (also referred to as code or instructions) that, when the program is run, causes the device to perform the method executed on the device side in any of the above embodiments.
[0239] This application also provides a readable storage medium storing a program (also referred to as code or instructions). When the program is run, it causes the device to perform the method executed on the device side in any of the above embodiments.
[0240] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0241] Those skilled in the art will understand that implementing all or part of the processes in the methods of the above embodiments can be accomplished by a program instructing related hardware. This program can be stored in a readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0242] In summary, the above description is merely an embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the disclosure of this application should be included within the scope of protection of this application.
Claims
1. A method for reading picture books, characterized in that, The method includes: Receive the first voice input for reading picture books; In response to the first voice input, a first image is acquired via a first camera, the first image containing the face of the first user who made the first voice input; A second image is acquired via a second camera in a first direction, the second image including a first picture book, and the audio of the first picture book is played. Receive a second voice input containing a wake word; In response to the second voice input, stop playing the audio of the first picture book, rotate to face the second user who made the second voice input, and interact with the second user via voice. A third image is acquired through the first camera in the second direction. The third image is an image containing the face of the first user, acquired through the first camera during the voice interaction with the second user. Receive third-party voice input for reading picture books; In response to the third voice input, rotate to the second direction; A fourth image is acquired via the first camera in the second direction; If the fourth image contains the face of the first user, and the angle between the second direction and the first direction is less than a threshold, rotate to the first direction; A fifth image is acquired via the second camera in the first direction; If the fifth image contains the first picture book, the audio of the first picture book continues to play.
2. The method according to claim 1, characterized in that, The method further includes: If the fifth image does not contain the first picture book, a first prompt message is output, which prompts the user to place the first picture book.
3. The method according to claim 1, characterized in that, The method further includes: If the fourth image contains the face of the first user, and the angle between the second direction and the first direction is greater than a threshold, a sixth image is acquired through the second camera in the second direction; If the sixth image contains the first picture book, the audio of the first picture book continues to play.
4. The method according to claim 3, characterized in that, The method further includes: If the sixth image does not contain the first picture book, a second prompt message is output, which prompts the user to place the first picture book.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: In the case that the fourth image does not contain the face of the first user, If a seventh image is acquired through the first camera in a third-party direction, and the seventh image contains the face of the first user, then an eighth image is acquired through the second camera in a third-party direction. If the eighth image contains the first picture book, the audio of the first picture book continues to play.
6. The method according to claim 5, characterized in that, The method further includes: If the eighth image does not contain the first picture book, a third prompt message is output, which prompts the user to place the first picture book.
7. The method according to any one of claims 1-4, characterized in that, The method further includes: In the case that the fourth image does not contain the face of the first user, If no image containing the first user's face is acquired through the first camera within the first time period, the device will enter a sleep state or a power-off state.
8. The method according to any one of claims 1-4, characterized in that, Before acquiring the second image via the second camera in the first direction, the method further includes: Turn off the first camera.
9. A method for reading picture books, characterized in that, The method includes: Receive the first voice input for reading picture books; In response to the first voice input, a first image is acquired through a first camera. The first image contains the face of the first user who made the first voice input, and the first image is acquired while facing a second direction. A second image is acquired via a second camera in a first direction, the second image including a first picture book, and the audio of the first picture book is played. Receive a second voice input containing a wake word; In response to the second voice input, the audio of the first picture book is stopped, the device rotates to face the second user who made the second voice input, and engages in voice interaction with the second user. During the voice interaction with the second user, the image captured by the first camera does not contain the face of the first user. Receive third-party voice input for reading picture books; In response to the third voice input, rotate to the second direction; A fourth image is acquired via the first camera in the second direction; If the fourth image contains the face of the first user, and the angle between the second direction and the first direction is less than a threshold, rotate to the first direction; A fifth image is acquired via the second camera in the first direction; If the fifth image contains the first picture book, the audio of the first picture book continues to play.
10. The method according to claim 9, characterized in that, The method further includes: If the fifth image does not contain the first picture book, a first prompt message is output, which prompts the user to place the first picture book.
11. The method according to claim 9, characterized in that, The method further includes: If the fourth image contains the face of the first user, and the angle between the second direction and the first direction is greater than a threshold, a sixth image is acquired through the second camera in the second direction; If the sixth image contains the first picture book, the audio of the first picture book continues to play.
12. The method according to claim 11, characterized in that, The method further includes: If the sixth image does not contain the first picture book, a second prompt message is output, which prompts the user to place the first picture book.
13. The method according to any one of claims 9-12, characterized in that, The method further includes: In the case that the fourth image does not contain the face of the first user, If a seventh image is acquired through the first camera in a third-party direction, and the seventh image contains the face of the first user, then an eighth image is acquired through the second camera in a third-party direction. If the eighth image contains the first picture book, the audio of the first picture book continues to play.
14. The method according to claim 13, characterized in that, The method further includes: If the eighth image does not contain the first picture book, a third prompt message is output, which prompts the user to place the first picture book.
15. The method according to any one of claims 9-12, characterized in that, The method further includes: In the case that the fourth image does not contain the face of the first user, If no image containing the first user's face is acquired through the first camera within the first time period, the device will enter a sleep state or a power-off state.
16. The method according to any one of claims 9-12, characterized in that, Before acquiring the second image via the second camera in the first direction, the method further includes: Turn off the first camera.
17. An electronic device, characterized in that, include: A memory, and one or more processors; the memory is coupled to the one or more processors, the memory being used to store program code, the program code including instructions, the one or more processors invoking the instructions to cause the electronic device to perform the method as described in any one of claims 1-16.
18. A readable storage medium comprising instructions, characterized in that, When the instructions are executed on the device, the device causes the device to perform the method as described in any one of claims 1-16.
Citation Information
Patent Citations
Reading monitoring method, reading robot and computing equipment
CN112084978A
Responding to a user query based on captured images and audio
US20230005471A1