Writing content acquisition method and device, equipment and storage medium

By acquiring and processing pre-shot images and writing images of book pages, and generating target images in real time, the problem of long waiting time for homework diagnosis in existing educational products is solved, thereby improving learning efficiency and feedback speed.

CN120877288APending Publication Date: 2025-10-31GUANGDONG XIAOTIANCAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410537649.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing educational products, users have to wait for diagnosis and grading after completing their assignments, which wastes time and leads to low learning efficiency.

Method used

By acquiring pre-shot images of the target book page and images of the book page written by the user, the user's writing content is generated in real time. Image processing technology is then used for registration and fusion to generate the target image.

Benefits of technology

It enables real-time feedback for homework correction and diagnosis, allowing users to immediately see the written content and errors, thus improving learning efficiency and feedback speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877288A_ABST
    Figure CN120877288A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a written content obtaining method and device, equipment and a storage medium, and the method comprises the steps: obtaining a pre-shot image comprising a target book page, obtaining a writing image comprising a book page written by a user under the condition that the user is in a writing state, the target book page comprises a book page written by the user, and obtaining a target image corresponding to the written content of the user according to the pre-shot image and the written image. By means of the method, the content written by the user can be obtained in real time, real-time feedback of homework correction, diagnosis and other functions is achieved, the user can immediately see the content written by the user and possible errors or improvement points, and therefore the learning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image generation technology, and to, but is not limited to, a method, apparatus, device, and storage medium for obtaining written content. Background Technology

[0002] Doing homework is not only an important way for students to consolidate their knowledge, but also a crucial method for assessing their understanding. Therefore, existing educational products typically offer auxiliary and diagnostic functions based on student homework scenarios. However, in most of these products, users have to actively click on grading or other functions after completing the homework, with homework diagnosis and assistance only occurring after this step. This requires a lengthy wait for the diagnostic process to complete, wasting a significant amount of the user's time.

[0003] Therefore, how to improve users' learning efficiency is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the method, apparatus, device, and storage medium for acquiring written content provided in this application embodiment can actively identify user responses and generate user work images in real time. The method, apparatus, device, and storage medium for acquiring written content provided in this application embodiment are implemented as follows:

[0005] This application provides a method for obtaining written content, including:

[0006] Acquire pre-taken images, including pages of the target book;

[0007] When the user is in a writing state, acquire a writing image including the book page written by the user, wherein the target book page includes the book page written by the user;

[0008] Based on the pre-shot image and the writing image, a target image corresponding to the user's writing content is obtained.

[0009] In some embodiments, acquiring a pre-captured image including a target book page includes:

[0010] If a preview image containing any page of a book is captured, determine whether the page of the book in the preview image is obscured;

[0011] If the book page in the preview image is not obstructed, the pre-shot image is obtained based on the preview image, and the target book page includes the book page in the preview image.

[0012] In some embodiments, acquiring a pre-captured image including a target book page includes:

[0013] If a preview image including any of the book pages is captured, determine whether the book corresponding to the book page in the preview image is in a stable state;

[0014] When the book corresponding to the book page in the preview image is in the stable state, the pre-shot image is obtained based on the preview image, and the target book page includes the book page in the preview image.

[0015] In some embodiments, obtaining the pre-shot image based on the preview image includes:

[0016] Obtain the book page area from the preview image;

[0017] The book page area is corrected to obtain the pre-shot image.

[0018] In some embodiments, obtaining the book page area in the preview image includes:

[0019] Identify multiple outline corner points of the book page in the preview image, and obtain the book page area based on the multiple outline corner points.

[0020] In some embodiments, obtaining the book page area in the preview image includes:

[0021] Detect the number of book pages in the preview image;

[0022] If there are multiple book pages in the preview image, determine whether the multiple book pages belong to the same book;

[0023] When the multiple book pages belong to the same book, obtain the multiple book page regions corresponding to the multiple book pages.

[0024] In some embodiments, the step of correcting the book page area to obtain the pre-shot image includes:

[0025] The correction process is performed on each of the plurality of book page regions to obtain at least one pre-shot image corresponding to the plurality of book page regions.

[0026] In some embodiments, acquiring a writing image including the book page written by the user while the user is in a writing state includes:

[0027] If a preview image including the user's body parts is captured, detect whether the user is in a writing state;

[0028] When the user is in the writing state, acquire a writing image including the book page written by the user.

[0029] In some embodiments, acquiring a writing image including the book page written by the user includes:

[0030] Obtain the book page area shown in the preview image where the user has written on the book page;

[0031] The area of ​​the book page written by the user is corrected to obtain a corrected image, which is the writing image.

[0032] In some embodiments, after correcting the book page area of ​​the user-written book page to obtain the corrected image, the method further includes:

[0033] The writing image is cropped to obtain a processed writing image, which does not include the user's body parts;

[0034] The processed writing image is fused with the pre-shot image to obtain the target image.

[0035] In some embodiments, the step of image fusion of the processed writing image and the pre-captured image to obtain the target image includes:

[0036] Feature extraction is performed on the processed writing image and the pre-shot image respectively to obtain feature points of the processed writing image and feature points of the pre-shot image;

[0037] Based on the feature points of the processed writing image and the feature points of the pre-shot image, the processed writing image and the pre-shot image are fused to obtain the target image.

[0038] In some embodiments, the step of image fusion of the processed writing image and the pre-shot image based on feature points of the processed writing image and feature points of the pre-shot image includes:

[0039] Feature point matching is performed on the feature points of the processed writing image and the feature points of the pre-shot image to obtain matched feature point pairs;

[0040] Based on the matched feature point pairs, the processed writing image and the pre-shot image are fused together.

[0041] In some embodiments, the step of image fusion of the processed writing image and the pre-shot image based on feature points of the processed writing image and feature points of the pre-shot image further includes:

[0042] Feature point matching is performed on the feature points of the processed writing image and the feature points of the pre-shot image to obtain a mapping matrix that maps the processed writing image to the pre-shot image;

[0043] Based on the mapping matrix, the processed writing image and the pre-shot image are fused together.

[0044] In some embodiments, the pre-captured image further includes a pre-captured question area. The step of performing feature extraction processing on the processed writing image and the pre-captured image respectively to obtain feature points of the processed writing image and feature points of the pre-captured image includes:

[0045] Identify the target question region in the processed writing image, perform feature extraction processing on the target question region, and obtain the feature points of the target question region;

[0046] Feature extraction is performed on the pre-shot question region in the pre-shot image to obtain the feature points of the pre-shot question region.

[0047] In some embodiments, the pre-captured image further includes a pre-captured answer region. The step of fusing the processed writing image with the pre-captured image based on feature points of the processed writing image and feature points of the pre-captured image to obtain the target image includes:

[0048] Identify the target answer region in the processed writing image;

[0049] Based on the feature points of the target question area and the feature points of the pre-shot question area, determine the mapping position of the target answer area in the pre-shot image;

[0050] The target answer region is mapped to the mapping position in the pre-shot image to obtain the target image.

[0051] In some embodiments, after obtaining the target image corresponding to the user's writing content based on the pre-captured image and the writing image, the method further includes:

[0052] If the user's preset action is detected, the target image is displayed.

[0053] This application provides an embodiment of a written content acquisition device, comprising:

[0054] The acquisition module is used to acquire pre-shot images, including pages of the target book;

[0055] The acquisition module is also used to acquire a writing image including the book page written by the user when the user is in a writing state, wherein the target book page includes the book page written by the user;

[0056] The generation module is used to obtain a target image corresponding to the user's writing content based on the pre-shot image and the writing image.

[0057] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.

[0058] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method described in this application embodiment.

[0059] This application provides a method, apparatus, computer device, and computer-readable storage medium for acquiring written content. By acquiring a pre-shot image including a target book page, and while the user is writing, acquiring a written image including the page the user is writing on, the target book page includes the page the user is writing on. Based on the pre-shot image and the written image, a target image corresponding to the user's writing content is obtained. This allows for real-time acquisition of the user's written content, enabling real-time feedback for functions such as homework correction and diagnosis. Users can immediately see their writing content and potential errors or areas for improvement, thereby improving learning efficiency and feedback speed, and solving the technical problems mentioned in the background art. Attached Figure Description

[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0061] Figure 1 This is an application scenario diagram of a method for obtaining written content disclosed in an embodiment of this application;

[0062] Figure 2 This is a schematic diagram illustrating the implementation process of a method for obtaining written content disclosed in an embodiment of this application;

[0063] Figure 3 This is a schematic diagram illustrating the implementation process of another method for obtaining written content disclosed in an embodiment of this application;

[0064] Figure 4This is an overall flowchart of a method for obtaining written content disclosed in an embodiment of this application;

[0065] Figure 5 This is a schematic diagram illustrating the implementation process of another method for obtaining written content disclosed in an embodiment of this application;

[0066] Figure 6 This is a screenshot illustrating a method for obtaining written content disclosed in an embodiment of this application;

[0067] Figure 7 This is an overall flowchart of another method for obtaining written content disclosed in an embodiment of this application;

[0068] Figure 8 This is a schematic diagram of the structure of a written content acquisition device disclosed in an embodiment of this application;

[0069] Figure 9 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0072] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0073] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0074] In view of this, embodiments of this application provide a method for obtaining written content, which is applied to smart electronic devices. Figure 1 This is an application scenario diagram of a method for obtaining written content, as illustrated in one embodiment. For example... Figure 1As shown, a user can carry, wear, or use electronic device 10, which may include, but is not limited to, learning machines, mobile phones, tablets, laptops, PCs (Personal Computers), etc. The functions implemented by this method can be achieved by the processor in the electronic device calling program code. Of course, the program code can be stored in the computer's storage medium. It can be seen that the electronic device includes at least a processor and a storage medium.

[0075] Figure 2 This is a schematic diagram illustrating the implementation flow of a method for obtaining written content provided in an embodiment of this application. For example... Figure 2 As shown, the method may include the following steps 201 to 203:

[0076] Step 201: Obtain a pre-shot image including the target book pages.

[0077] In this embodiment, a camera or scanning device is used to photograph or scan the target book page to obtain a clear pre-captured image. Automatic focus and light adjustment functions are used to ensure the image quality of the pre-captured image. Optionally, background interference can be removed during image acquisition to highlight the content of the book page.

[0078] Step 202: When the user is in a writing state, acquire a writing image including the book page written by the user, wherein the target book page includes the book page written by the user.

[0079] In this embodiment, while the user is writing, an image of the page being written on is captured using the same camera or scanning device. The dynamic changes during the writing process are captured in real time, and multiple images are stitched together to obtain a complete writing image. Optionally, the camera parameters can be adjusted as needed for different writing tools and methods to ensure a clear writing image is captured.

[0080] Step 203: Based on the pre-shot image and the writing image, obtain the target image corresponding to the user's writing content.

[0081] In this embodiment, firstly, image processing techniques are used to preprocess the pre-captured image and the written image, such as denoising and contrast enhancement, to improve the accuracy of subsequent processing. Next, image registration techniques are used to register the written image with the pre-captured image, finding the correspondence between them. After registration, the content written by the user is accurately mapped to the corresponding position on the target book page based on the correspondence, thus obtaining the target image corresponding to the user's writing content. Finally, the obtained target image is post-processed, such as removing overlapping parts and smoothing edges, to improve the image's clarity and aesthetics. Optionally, during image registration, the curvature and deformation of the book page need to be considered, and a non-rigid transformation model is used for registration. For example, feature matching algorithms, such as SIFT (Scale-Invariant Feature Transform) or SURF (Speeded-Up Robust Features), can be used to improve the accuracy and robustness of registration. For complex writing content and page layouts, deep learning methods, such as CNN (Convolutional Neural Network) or RNN (Recurrent Neural Network), can be used to achieve more accurate registration and mapping. After obtaining the target image, image inpainting can be performed to fill in any missing or damaged areas, improving the image's integrity and readability. For example, image enhancement techniques can be used, such as removing text backgrounds and enhancing text clarity, to improve the image's visual quality and visualization. Finally, the generated target image is evaluated and verified to ensure its consistency and accuracy with the original book pages.

[0082] Through the embodiments of this application, by combining pre-captured images and writing images, a target image corresponding to the user's writing content can be generated more accurately. The pre-captured image provides basic structural and layout information of the target book page, while the writing image provides the actual content written by the user. Combining the two generates a more complete and accurate target image. This enables real-time feedback for functions such as homework grading and diagnostics, thereby improving learning efficiency.

[0083] In the above Figure 2 As shown in the illustration, this application embodiment also provides a schematic diagram of the implementation flow of a method for obtaining written content. For example... Figure 3 As shown, the method may include the following steps 301 to 304:

[0084] Step 301: If a preview image including any book page is acquired, determine whether the book page in the preview image is obscured.

[0085] In this embodiment, firstly, a preview image including any page of a book is acquired. Then, real-time image processing technology, such as a real-time object detection algorithm, is used to continuously monitor whether occlusion occurs in the image. Deep learning models, such as YOLO (You Only Look Once) or SSD (Single Shot MultiBoxDetector), can be used to perform real-time occlusion detection on the image. Detected occlusions are categorized, such as hand occlusion or book edge occlusion. Based on machine learning algorithms, the model is trained to identify different types of occlusion and take corresponding processing measures. When the system detects an occlusion, it promptly issues a warning to the user so that the user can adjust their posture or remove the occluder. Feedback to the user can be provided through sound prompts, vibrations, or prompt icons on the graphical interface.

[0086] Step 302: If the book page in the preview image is not obstructed, obtain the pre-shot image based on the preview image, wherein the target book page includes the book page in the preview image.

[0087] In this embodiment, when the book pages in the preview image are unobstructed, automatic focus and exposure control can be performed to ensure appropriate image sharpness and brightness. This can be achieved using the camera's autofocus and auto-exposure functions, or through real-time adjustments via software algorithms. After capturing multiple preview images consecutively, image fusion techniques (such as mean overlay or multi-frame fusion) are used to generate a high-quality pre-shot image. By fusing images from multiple angles or exposure levels, noise can be reduced and image sharpness increased.

[0088] Step 303: When the user is in a writing state, acquire a writing image including the book page written by the user, wherein the target book page includes the book page written by the user.

[0089] Step 303 is implemented in the same way as step 202 above, and will not be described again in this application.

[0090] Step 304: Based on the pre-shot image and the writing image, obtain a target image corresponding to the user's writing content.

[0091] Step 304 is implemented in the same way as step 203 above, and will not be described again in this application.

[0092] In this embodiment, by selecting an unobstructed preview image for pre-capture, it is possible to ensure that the target image matches the content actually written by the user. A clear target image helps to accurately identify the written content and ensures that no information is omitted or distorted. By reducing the processing of invalid images, the system can focus more on processing valid images, thereby improving processing efficiency and response speed.

[0093] As an example, acquiring a pre-captured image including a target book page includes: when a preview image including any of the book pages is acquired, determining whether the book corresponding to the book page in the preview image is in a stable state.

[0094] Specifically, the device can integrate sensors, such as accelerometers or gyroscopes, to monitor whether the book is stationary. By detecting the book's movement, the system can determine whether the book is in a stable state. Alternatively, multiple preview images can be acquired, and the stability of the book can be determined by analyzing changes in the position and shape of the pages in the preview images. For example, if the position of the pages in several consecutive preview images does not change significantly, it may indicate that the book is in a stable state. Furthermore, a threshold for determining a stable state can be set, such as the maximum allowable range of page position changes or the standard deviation of position changes in consecutive preview images. If the position change of the pages is below the set threshold, the book is confirmed to be in a stable state. The device can also monitor the book's movement in real time and immediately acquire pre-captured images when a stable state is detected. This ensures that the acquired pre-captured images are taken when the book is in a stable state.

[0095] Furthermore, if the book corresponding to the book page in the preview image is in the stable state, the pre-shot image is obtained based on the preview image, and the target book page includes the book page in the preview image.

[0096] Specifically, once the book is confirmed to be in a stable state, pre-capture images are immediately acquired to ensure high image quality and accurate correspondence with the actual written content. Optionally, a high-resolution camera or image acquisition device is prioritized to acquire pre-capture images, improving image clarity and detail capture. Automatic focus and exposure control can also be implemented to ensure appropriate image clarity and brightness. This can be achieved using the camera's autofocus and auto-exposure functions, or through real-time adjustments via software algorithms. After capturing multiple preview images consecutively, image fusion techniques (such as mean overlay or multi-frame fusion) are used to generate high-quality pre-capture images. By fusing images from multiple angles or exposure levels, noise can be reduced and image clarity increased.

[0097] In this embodiment, by determining whether the book is in a stable state, pre-capture images can be avoided when the book is shaking or moving, thereby reducing image blurring or distortion caused by motion. This results in clearer and more accurate pre-capture images. Pre-capture images acquired in a stable state are of high quality, reducing the complexity of subsequent image processing. The clearer and more accurate pre-capture images acquired in a stable state are beneficial for subsequently obtaining the target image corresponding to the user's writing content based on the pre-capture image and the writing image. This improves the accuracy and reliability of the target image and reduces recognition errors caused by poor image quality.

[0098] As an example, obtaining the pre-shot image based on the preview image includes: obtaining the book page area of ​​the book page in the preview image.

[0099] Specifically, image processing algorithms, such as edge detection or template matching, are used to detect book page regions in the preview image. Furthermore, a model can be pre-trained to recognize the shape, boundaries, or features of the book pages to accurately detect their location.

[0100] As an example, obtaining the book page region in the preview image includes: identifying multiple contour corner points of the book page in the preview image, and obtaining the book page region based on the multiple contour corner points.

[0101] Specifically, image processing techniques, such as edge detection or color segmentation, are used to extract the main outline and edge features of the book page. Image enhancement is then performed to improve the visibility of the book page's edges and corners. Edge detection algorithms, such as Canny edge detection, are used to find the main outline of the book page. Corner detection algorithms, such as Harris corner detection or Shi-Tomasi corner detection, are used to identify the corners of the book page. The detected corners are then filtered to remove irrelevant or noisy points. Based on the corner position information, the corners are connected to construct the book page boundary. Based on the connected corners, the boundary region of the book page is determined. Convex hull algorithms or minimum bounding rectangle algorithms can be used to expand the area enclosed by the corners into a complete book page region. The book page region is then corrected, including perspective transformation or affine transformation, to eliminate any possible distortions and artifacts. The correction process can adjust the image projection based on the book page's feature points, such as corner positions and boundary lines, to obtain an accurate book page image.

[0102] This application's embodiments can more accurately determine the boundaries of book pages by identifying multiple contour corner points, avoiding errors that may be caused by a single feature point or edge. Increasing the number of contour corner points can improve the accuracy of boundary detection, especially for book pages with complex shapes or deformations.

[0103] As an example, obtaining the book page region in the preview image includes: detecting the number of book pages in the preview image.

[0104] Specifically, image segmentation techniques, such as semantic segmentation algorithms based on pixel color, texture, or deep learning, are used to separate the book pages in the preview image from the background, thus obtaining the independent region of each book page. Edge detection algorithms, such as Canny edge detection, can be used to find the contours of the book pages in the image, and then the number of contours can be calculated.

[0105] Furthermore, if there are multiple book pages in the preview image, it is determined whether the multiple book pages belong to the same book.

[0106] Specifically, feature extraction and matching techniques, such as SIFT or SURF algorithms, are used to extract feature descriptors from book pages and match them to determine if they come from the same book. Geometric transformations, such as perspective or affine transformations, are used to correct the detected book pages to give them a consistent viewpoint and scale. Then, overlapping areas or aligned positional relationships are used to determine if they belong to the same book.

[0107] Furthermore, if the multiple book pages belong to the same book, multiple book page regions corresponding to the multiple book pages are obtained.

[0108] Specifically, for multiple pages belonging to the same book, their positions in the original image can be repositioned or cropped based on their location information or geometric relationships, forming a collection of multiple book page regions.

[0109] This application embodiment detects all book pages in the preview image, ensuring that no page is missed, thus guaranteeing complete capture of the book content. It avoids misidentifying other objects or non-book pages as book pages, improving the accuracy and reliability of recognition. It also avoids mixing pages from different books together, reducing potential errors or confusion in subsequent processing. Furthermore, it ensures that only pages from the same book are identified as a group, guaranteeing consistency and accuracy in subsequent processing.

[0110] As an example, the step of correcting the book page area to obtain the pre-shot image includes: performing the correction process on each of the plurality of book page areas to obtain at least one pre-shot image corresponding to the plurality of book page areas.

[0111] Specifically, image processing techniques such as Canny edge detection or contour detection algorithms are used to find the boundaries of each book page region. For each detected book page region, perspective transformation is applied to correct its angles and distortions to ensure that its boundaries are horizontal and vertical. Each book page region is resized according to a preset standard size to ensure they have similar dimensions and proportions. Images corresponding to each corrected book page region are extracted from the original pre-captured images. Enhancement processing, such as noise reduction and contrast enhancement, is then applied to the extracted images of each book page region to improve image quality and sharpness.

[0112] Furthermore, the book page area is corrected to obtain the pre-shot image.

[0113] Specifically, if perspective distortion exists on the book pages in the preview image, it can be eliminated using perspective correction algorithms, such as perspective transformation or projection transformation. Image rotation, scaling, or cropping can also be performed to position and size the book pages appropriately in the preview image. When performing correction, boundary conditions must be considered to ensure the corrected image is not truncated or distorted. Boundary extension or padding techniques can be used to handle boundary conditions, preserving all important information in the image and avoiding boundary effects. The parameters of the correction algorithm, such as rotation angle and cropping size, should be adjusted to optimize the quality of the corrected image. The optimal parameter settings can be determined through experimentation and evaluation of different parameter combinations. The corrected image should be visualized or quantitatively evaluated to verify that the correction effect meets expectations.

[0114] The accuracy and stability of the correction effect can be evaluated by comparing the image features before and after correction or by performing target detection.

[0115] In this embodiment, by calibrating the geometry and structure in the preview image, any form of distortion and falsification caused by shooting angle, lens distortion, or perspective effects is eliminated. This ensures that the shape and proportion of the book pages in the image are accurate. The calibrated preview image is easier to automatically identify, segment, and analyze. By eliminating distortion and optimizing image quality, the calibration process improves the accuracy and robustness of subsequent recognition and analysis algorithms.

[0116] This application also provides an overall flowchart of a method for obtaining written content, such as... Figure 4 As shown in the embodiment of this application, the stability of the book is determined by the book region detection algorithm and the book stability detection algorithm. If the book is stable, the next step is performed; otherwise, the next step is not performed.

[0117] Image processing techniques, such as edge detection and color segmentation, are used to detect potential book regions within an image. This can be achieved by detecting features such as edges, shape, or color. A lightweight object detection algorithm is employed to locate book pages within an image.

[0118] Once a book region is detected, image stabilization techniques, such as image registration or visual inertial navigation, can be used to ensure the book remains stable in the image for subsequent processing. Based on the book region detection results across multiple frames, the Intersection over Union (IoU) of the book bounding boxes is used to determine whether the book is stable.

[0119] Within a stable book area, object detection algorithms, such as deep learning-based object detection models, are used to detect items that may obscure the book's content. An occlusion detection algorithm is then used to determine if occlusion exists; only if no occlusion is found does the next step proceed.

[0120] Under stable and unobstructed conditions within the book area, image matching or feature matching techniques are used to detect whether multiple pages belong to the same book, allowing for merging or separation. It also determines which pages belong to the same book as the main pages. A lightweight keypoint detection algorithm is used to detect the four corner points of the book's outline.

[0121] For images of the book area, geometric transformation techniques, such as perspective transformation, are used to correct the book area into a rectangle or square for subsequent processing and analysis.

[0122] For text images within the book area, which may be curved or distorted, image processing techniques such as curve fitting or perspective transformation are used to correct the text image to improve the accuracy of text recognition.

[0123] Within the corrected book area, techniques such as object detection or bounding box regression are used to detect the position and boundary of the question boxes, and to detect the position of all questions within the pre-shot image for subsequent identification and analysis.

[0124] Similarly, the location and boundaries of possible answer boxes within the book area are detected, and the location of all answer areas within the pre-shot image is detected in order to identify and analyze students' answers.

[0125] This application's embodiments eliminate any form of distortion and falsification caused by shooting angle, lens distortion, or perspective effects by calibrating the geometry and structure in the preview image, ensuring the accurate shape and proportion of the book pages in the image. The calibrated preview image is easier to automatically recognize, segment, and analyze, thereby improving the accuracy and robustness of subsequent recognition and analysis algorithms. Text images within the book area are corrected to improve the accuracy of text recognition. The position and boundaries of the question and answer boxes are detected, enabling rapid and accurate location and recognition of students' answers.

[0126] In the above Figure 2 As shown in the illustration, this application embodiment also provides a schematic diagram of the implementation flow of a method for obtaining written content. For example... Figure 5 As shown, the method may include the following steps 501 to 504:

[0127] Step 501: Obtain a pre-shot image including the target book pages.

[0128] Step 501 is implemented in the same way as step 201 above, and will not be described again in this application.

[0129] Step 502: If a preview image including the user's body parts is acquired, detect whether the user is in a writing state.

[0130] In this embodiment, after capturing preview images including parts of the user's body using a camera or other image acquisition device, computer vision technology is used to detect the user's body posture and hand movements in the preview images. Specifically, deep learning models or traditional image processing algorithms can be used to detect specific gestures or movements to determine whether the user is in a writing state. A model can also be trained to recognize writing actions, such as the gesture of holding a pen and moving it across paper.

[0131] Step 503: When the user is in the writing state, acquire a writing image including the book page written by the user.

[0132] In this embodiment, if a user is detected to be writing, a writing image containing the book page written by the user is extracted from the preview image. Image segmentation techniques can be used to separate the writing area from the preview image. Further image processing, such as noise reduction and contrast enhancement, may be required for the acquired writing image to improve image quality and clarity. Illumination correction and edge detection can also be performed to ensure high-quality extracted writing images suitable for subsequent analysis and processing.

[0133] Step 504: Based on the pre-shot image and the writing image, obtain the target image corresponding to the user's writing content.

[0134] Step 504 is implemented in the same way as step 203 above, and will not be described again in this application.

[0135] This application embodiment can automatically detect whether a user is in a writing state by acquiring preview images of the user's body parts. When the user is confirmed to be writing, the system can promptly acquire an image of the page the user is writing on. This ensures that no writing content is missed, effectively capturing the user's writing process and content.

[0136] As an example, obtaining a writing image including the user's handwritten book page includes: obtaining the book page area of ​​the user's handwritten book page in a preview image.

[0137] Specifically, the first step is to identify the area of ​​the book page where the user is writing from the preview image. Image processing and computer vision techniques, such as object detection and segmentation, are then used to determine the location and boundaries of the book page in the preview image.

[0138] Furthermore, the area of ​​the book page written by the user is corrected to obtain a corrected image, which is the writing image.

[0139] Specifically, the identified book page areas undergo correction processing. This corrects distortion, perspective transformation, or rotation to ensure the content of the book pages appears more regular and accurate in the image. The corrected book page areas are then composited or extracted into a new image, which is the corrected writing image. This involves image stitching, cropping, or other correction operations to ensure the content of the book pages is displayed optimally. The final output is the corrected writing image, which can be used for further analysis, recognition, or processing to obtain the content or features of the user's handwriting.

[0140] This application embodiment obtains a region of the book page in a preview image where the user is writing, and then performs correction processing on that region to obtain a corrected image. This ensures that the obtained writing image more accurately matches the user's actual writing behavior. It improves the quality and accuracy of the writing image, thereby enhancing the effectiveness of subsequent processing and recognition, making the method for obtaining the writing content more reliable and effective.

[0141] As an example, after correcting the area of ​​the book page written by the user to obtain the corrected image, the method further includes: taking a screenshot of the writing image to obtain a processed writing image, wherein the processed writing image does not include the user's body parts.

[0142] Specifically, image processing techniques are used to identify the position and boundaries of the book pages in the preview image. Algorithms such as image segmentation or edge detection are used to determine the area of ​​the book page. Geometric transformation techniques such as perspective transformation or affine transformation are used to correct the book page area, deforming or rotating it to a standard rectangular shape. This ensures that the corrected image maintains the integrity and accuracy of the written content. In the corrected image, techniques such as cropping or masking are used to remove content other than the user's body parts, retaining only the image of the writing area. This ensures that the captured writing image is clearly visible and does not contain any distracting or irrelevant parts.

[0143] For example, this application provides a method for taking screenshots of written images, such as... Figure 6 As shown, Figure 6 This is a screenshot illustrating a method for obtaining written content provided in this application.

[0144] The image without hands is extracted from the original image, and only the handwritten answers in this part are displayed. The extraction rules are as follows: the large rectangle represents the book, and the two small squares represent the hands. Based on the relative positions of the hands and the book, four types are categorized: bottom type, top type, top and bottom divided into left and right types, and top and bottom and left and right of the same type. In the top type, the hands are all above the center line of the book. In this case, the minimum y-axis coordinate of the hand detection box is taken as the dividing line, such as... Figure 6 The dotted line in the upper middle is the dividing line, and the part below the dividing line is the screenshot area, as shown in the area where the ellipse in Figure 6 is located.

[0145] The lower part of the image shows the hands positioned below the center line of the book. In this case, the maximum value in the y-axis coordinate of the hand detection box is taken as the dividing line. Figure 6 The dotted line at the bottom center is the dividing line; the area above the dividing line is the screenshot area. Figure 6 The region where the ellipse is located.

[0146] The image is divided into two hands, one above and one below, one left and one right. The maximum x-coordinate of the left hand detection box, the minimum y-coordinate of the top hand detection box, the minimum x-coordinate of the right hand detection box, and the maximum y-coordinate of the bottom hand detection box are then used to determine the rectangle formed by these four coordinates. Figure 6 The rectangle formed by the dotted lines at the top, bottom, left, and right is the screenshot area. Figure 6 The region where the ellipse is located.

[0147] The top and bottom positions represent two hands, one above the other, and both hands simultaneously on the left or right. If both hands are on the left, the larger x-axis coordinate of the two hands is used as the dividing line, and the area to the right of this dividing line is the screenshot area. If both hands are on the right, the smaller x-axis coordinate of the two hands is used as the dividing line, and the area to the left of this dividing line is the screenshot area.

[0148] Furthermore, the processed writing image is fused with the pre-shot image to obtain the target image.

[0149] Specifically, the screenshot of the written image is blended with the pre-taken image using methods such as transparency overlay or image overlay to perfectly integrate the written content with the original book page. The blended target image should clearly display the user's written content while preserving the layout and format of the original page.

[0150] This application embodiment corrects the user's handwritten page to obtain a corrected image. Then, it performs a screenshot of the corrected image to remove the user's body parts. Finally, it fuses the processed writing image with a pre-captured image to obtain the target image. This effectively removes interference from the user's body parts, making the target image clearer and more focused on the written content, thus improving image quality and usability.

[0151] As an example, the step of fusing the processed writing image with the pre-shot image to obtain the target image includes: performing feature extraction processing on the processed writing image and the pre-shot image respectively to obtain feature points of the processed writing image and feature points of the pre-shot image.

[0152] Specifically, feature extraction algorithms (such as SIFT and SURF) are used to extract feature points and their descriptors from the processed writing image and the pre-captured image. Feature matching algorithms (such as nearest neighbor matching algorithms) are then used to match the feature points in the writing image and the pre-captured image to establish the correspondence between them.

[0153] Furthermore, based on the feature points of the processed writing image and the feature points of the pre-shot image, the processed writing image and the pre-shot image are fused to obtain the target image.

[0154] Specifically, by utilizing the correspondence of feature points, image registration algorithms such as RANSAC (Random Sample Consensus) and Homography (Homography Transformation) are used to register the written image so that it is aligned with the pre-shot image.

[0155] The registration process corrects for rotation, translation, and scaling in the written image to align it with the pre-captured image. Image fusion algorithms (such as image overlay or blending) are then used to merge the registered written image with the pre-captured image. Factors such as transparency and color matching may need to be considered during the fusion process to ensure the merged image looks natural and complete. The merged image is then output as the target image, containing both the user's handwriting and the book pages from the background pre-captured image.

[0156] This application's embodiments, through feature extraction processing, can more accurately match corresponding feature points between the written image and the pre-captured image, thereby improving the accuracy of image fusion. Feature extraction processing helps reduce image distortion and deformation during the fusion process, making the target image closer to the actual written content, thus improving readability and clarity.

[0157] As an example, the step of image fusion of the processed writing image and the pre-shot image based on the feature points of the processed writing image and the feature points of the pre-shot image includes: performing feature point matching on the feature points of the processed writing image and the feature points of the pre-shot image to obtain a matched feature point pair.

[0158] Specifically, feature extraction is performed on the processed writing image and the pre-captured image. Classic feature extraction algorithms such as SIFT, SURF, and ORB (Oriented Fast and Rotated BRIEF) can be used. The extracted feature points should possess stability, discriminative power, and reliability to ensure the accuracy of subsequent matching. Feature descriptors are used to describe the extracted feature points. Feature matching algorithms (such as nearest neighbor-based matching algorithms) are used to match the feature points of the writing image and the pre-captured image. Filtering mechanisms, such as the RANSAC (Random Sample Consensus) algorithm, can be employed to remove incorrect matches and improve the accuracy and stability of the matching process.

[0159] Furthermore, based on the matched feature point pairs, the processed writing image and the pre-shot image are fused together.

[0160] Specifically, based on the matched feature point pairs, image transformation is performed to align the written image onto the pre-captured image. Geometric transformations such as affine transformations and perspective transformations can be used. Interpolation algorithms (such as bilinear interpolation and bicubic interpolation) are then used to perform pixel-level fusion of the transformed written image to ensure seamless integration with the pre-captured image. Post-processing of the image can be performed as needed, such as adjusting brightness, contrast, and color balance, to ensure optimal quality of the fused image. Image enhancement techniques, such as sharpening and noise reduction, can be used to further improve image clarity and quality.

[0161] This application's embodiments achieve effective merging of the processed written image and the pre-captured image through feature point matching and image fusion, thereby obtaining a more accurate and clearer target image. This can improve the accuracy of recognizing and acquiring written content, providing a better user experience.

[0162] As an example, the step of image fusion of the processed writing image and the pre-shot image based on the feature points of the processed writing image and the feature points of the pre-shot image includes: performing feature point matching on the feature points of the processed writing image and the feature points of the pre-shot image to obtain a mapping matrix that maps the processed writing image to the pre-shot image.

[0163] Specifically, feature point detection algorithms, such as SIFT or SURF, are applied to both the processed writing image and the pre-captured image to detect key feature points and calculate descriptors for these feature points. Using the detected feature points and their descriptors, feature point matching is performed between the two images to find corresponding feature point pairs. This can be achieved using matching algorithms such as nearest neighbor matching or feature descriptor-based matching. Using the matched feature point pairs, the least squares method or other optimization algorithms are used to calculate a mapping matrix that maps the processed writing image to the pre-captured image. This mapping matrix describes the geometric transformation relationship between the two images and is typically an affine transformation matrix or a perspective transformation matrix.

[0164] Furthermore, based on the mapping matrix, the processed writing image and the pre-shot image are fused together.

[0165] Specifically, the processed writing image is fused with the pre-captured image using the calculated mapping matrix. Image interpolation techniques, such as bilinear interpolation or triangular interpolation, can be used to map pixels from the writing image onto the pre-captured image, followed by pixel fusion to generate the final target image.

[0166] This application embodiment achieves the process of mapping the written image to the pre-shot image by matching feature points between the processed written image and the pre-shot image to obtain a mapping matrix. This helps to accurately insert the user's written content into the target image, ensuring that the final target image retains the user's written pages and blends naturally with the original pre-shot image visually, without producing obvious traces or incongruous areas. It can improve the user experience, making the final generated target image more in line with user expectations, and possessing higher usability and readability.

[0167] As an example, the pre-captured image also includes a pre-captured question region. The step of performing feature extraction processing on the processed writing image and the pre-captured image respectively to obtain feature points of the processed writing image and feature points of the pre-captured image includes: identifying the target question region in the processed writing image, performing feature extraction processing on the target question region to obtain feature points of the target question region.

[0168] Specifically, image processing techniques (such as edge detection and image segmentation) are used to identify target title regions in written images. For example, text detection and recognition algorithms can be used to find title or table of contents regions in written images. Feature extraction is then performed on the identified target title regions. Common methods include local feature descriptors (such as SIFT, SURF, and ORB) or deep learning features (such as features extracted by CNNs). The extracted feature points should have the ability to describe key information within the region, such as corners and edges.

[0169] Furthermore, feature extraction is performed on the pre-shot question region in the pre-shot image to obtain the feature points of the pre-shot question region.

[0170] Specifically, feature extraction is performed on the pre-captured question region in the pre-captured image using the same method as for the written image, which will not be elaborated upon here. It is ensured that the feature points of the pre-captured question region can describe the key information of that region. Feature point matching algorithms (such as nearest neighbor matching or RANSAC algorithm) are used to match the feature points of the target question region in the written image with the feature points of the pre-captured question region in the pre-captured image.

[0171] The goal of matching is to find similar feature point pairs in the written image and the pre-shot image.

[0172] The feature extraction process in this application can help identify the target question region in the written image and the pre-shot question region in the pre-shot image, and extract their feature points. This can be used to match and fuse the written image and the pre-shot image more quickly, thereby improving the accuracy and stability of obtaining the target image.

[0173] As an example, the pre-captured image also includes a pre-captured answer region. The step of fusing the processed writing image with the pre-captured image based on the feature points of the processed writing image and the feature points of the pre-captured image to obtain the target image includes: identifying the target answer region in the processed writing image.

[0174] Specifically, image processing techniques, such as edge detection and color segmentation, are used to identify the target answer region in the processed written image. Preprocessing can be performed first, such as removing background noise or adjusting image brightness and contrast. Pattern recognition or machine learning algorithms, such as convolutional neural networks (CNNs), are then employed to accurately identify the target answer region.

[0175] Furthermore, based on the feature points of the target question area and the feature points of the pre-shot question area, the mapping position of the target answer area in the pre-shot image is determined.

[0176] Specifically, feature points of the target answer region are extracted from the processed writing image. These feature points can be corner points, edge points, or other salient feature points. Similarly, corresponding feature points are extracted from the pre-captured question region of the pre-captured image. Feature point matching algorithms, such as nearest neighbor matching or RANSAC, are used to match the feature points in the processed writing image with those in the pre-captured image. The relationship between the matching point pairs is determined for subsequent image fusion. Based on the positional relationship of the matching point pairs, the mapping position of the target answer region in the pre-captured image is calculated. This is achieved through geometric transformations, such as affine transformations or perspective transformations. The mapping position can be determined by fitting a transformation model or by using interpolation methods to obtain more accurate results.

[0177] Furthermore, the target answer region is mapped to the mapping position in the pre-shot image to obtain the target image.

[0178] Specifically, the target answer region is extracted from the processed written image and mapped to its corresponding location in the pre-captured image. Image overlay or blending techniques can be used to fuse the target answer region with the pre-captured image to generate the final target image.

[0179] This application embodiment achieves image fusion based on the processed written image and the pre-captured image to identify the target answer region and map it to the mapped position in the pre-captured image. This enables accurate acquisition of the written content, including the written answer, thereby achieving the recognition and analysis of the written content.

[0180] As an example, after obtaining the target image corresponding to the user's writing content based on the pre-shot image and the writing image, the method further includes: displaying the target image when a preset operation of the user is detected.

[0181] Specifically, a dedicated module can be incorporated into the device to detect preset user actions. This module can use computer vision or sensor technology to capture user movements or gestures. It continuously monitors user behavior to identify specific preset actions, such as clicking a specific button, a gesture, or a voice command. It then determines the conditions that trigger the display of the target image, including button clicks, gestures, and specific words. Once the preset action is detected, the device triggers a corresponding response mechanism. After triggering the response mechanism, the system displays the target image generated by fusing the processed written image with the pre-captured image to the user. This is achieved by displaying the image on the screen or presenting it within the application interface. The device also provides feedback to inform the user that their action has been successfully recognized and the target image has been displayed.

[0182] Through the embodiments of this application, users can conveniently view their written content when needed without manually triggering the display. This automated display method improves the user experience.

[0183] This application also provides an overall flowchart of a method for obtaining written content, such as... Figure 7 As shown in this embodiment, the stability of the book is determined by a book region detection algorithm and a book stability detection algorithm. For example, a lightweight object detection algorithm can be used to detect the book region in the image. Image stabilization techniques (such as image registration, visual inertial navigation, etc.) are used to ensure the stability of the book region. If stable, proceed to the next step; otherwise, do not proceed to the next step.

[0184] Once the book is stable, the first step is to use hand detection and pen-holding detection algorithms. A deep learning object detection model can be used to detect the hand region in the image. Hand pose estimation techniques are then used to detect the hand's pen-holding state.

[0185] Then, it determines which page the user is currently answering on, i.e., the main page. Next, it performs book detection on the main page, obtaining the four corner points of the book. Specifically, a lightweight object detection algorithm is used to detect the main area and other potentially interfering objects. It is then determined whether the main area contains the book; if so, the next step is performed.

[0186] Finally, based on the book detection results, trapezoidal and curvature corrections are performed on the main pages. For the detected book areas, perspective transformation techniques are applied to correct them into rectangles or squares. For text images, curve fitting or perspective transformation is used for curvature correction.

[0187] Based on user needs or specified areas, extract the top part of the book and the handwritten answer section.

[0188] For the extracted region, feature points are extracted using feature point extraction algorithms (such as SIFT, SURF, etc.).

[0189] Use feature point matching algorithms (such as the FLANN matcher) to match feature points between different images.

[0190] Based on the results of feature point matching, image fusion is performed to combine the features of different images.

[0191] When the user taps the screen, the merged image is displayed.

[0192] This application's embodiments utilize advanced technologies such as deep learning object detection models and feature point matching algorithms to improve the accurate detection and recognition capabilities of book areas and written content, thereby enhancing the overall system accuracy. Image stabilization technology ensures the stability of the book area, reducing recognition errors caused by image jitter or instability and improving system stability and reliability. By correcting and curvature in the book area, as well as image fusion, the user experience of acquiring written content is improved, resulting in clearer and more accurate information, thus enhancing the user experience.

[0193] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0194] Based on the foregoing embodiments, this application provides a device for acquiring written content. The device includes various modules and units included in each module, which can be implemented by a processor; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc.

[0195] Figure 8 This is a schematic diagram of the structure of a device for acquiring written content provided in an embodiment of this application, as shown below. Figure 8 As shown, the device 800 includes an acquisition module 801 and a generation module 802, wherein:

[0196] The acquisition module 801 is used to acquire pre-shot images including pages of the target book;

[0197] The acquisition module 801 is further configured to acquire a writing image including the book page written by the user when the user is in a writing state, wherein the target book page includes the book page written by the user;

[0198] The generation module 802 is used to obtain a target image corresponding to the user's writing content based on the pre-shot image and the writing image.

[0199] In some embodiments, the acquisition module 801 is further configured to determine whether the book page in the preview image is obstructed when a preview image including any book page is acquired;

[0200] The acquisition module 801 is further configured to acquire the pre-shot image based on the preview image when there is no obstruction to the book page in the preview image, wherein the target book page includes the book page in the preview image.

[0201] In some embodiments, the acquisition module 801 is further configured to determine whether the book corresponding to the book page in the preview image is in a stable state when a preview image including any of the book pages is acquired.

[0202] The acquisition module 801 is further configured to acquire the pre-shot image based on the preview image when the book corresponding to the book page in the preview image is in the stable state, wherein the target book page includes the book page in the preview image.

[0203] In some embodiments, the acquisition module 801 is further configured to acquire the book page area of ​​the book page in the preview image;

[0204] The generation module 802 is also used to perform correction processing on the book page area to obtain the pre-shot image.

[0205] In some embodiments, the acquisition module 801 is further configured to identify multiple outline corner points of the book page in the preview image, and acquire the book page area based on the multiple outline corner points.

[0206] In some embodiments, the acquisition module 801 is further configured to detect the number of book pages in the preview image;

[0207] The acquisition module 801 is also used to determine whether the multiple book pages in the preview image belong to the same book when there are multiple book pages in the preview image.

[0208] The acquisition module 801 is further configured to acquire multiple book page regions corresponding to the multiple book pages when the multiple book pages belong to the same book.

[0209] In some embodiments, the generation module 802 is further configured to perform the correction processing on each of the plurality of book page regions to obtain at least one pre-shot image corresponding to the plurality of book page regions.

[0210] In some embodiments, the acquisition module 801 is further configured to detect whether the user is in a writing state when a preview image including the user's body parts is acquired;

[0211] The acquisition module 801 is further configured to acquire a writing image including the book page written by the user when the user is in the writing state.

[0212] In some embodiments, the acquisition module 801 is further configured to acquire the book page area of ​​the book page written by the user in the preview image;

[0213] The generation module 802 is also used to perform correction processing on the book page area of ​​the book page written by the user to obtain a corrected image, wherein the corrected image is the writing image.

[0214] In some embodiments, the generation module 802 is further configured to perform screenshot processing on the writing image to obtain a processed writing image, wherein the processed writing image does not include the user's body parts;

[0215] The processed writing image is fused with the pre-shot image to obtain the target image.

[0216] In some embodiments, the acquisition module 801 is further configured to perform feature extraction processing on the processed writing image and the pre-shot image respectively, to obtain feature points of the processed writing image and feature points of the pre-shot image;

[0217] The generation module 802 is further configured to perform image fusion between the processed writing image and the pre-shot image based on the feature points of the processed writing image and the feature points of the pre-shot image to obtain the target image.

[0218] In some embodiments, the acquisition module 801 is further configured to perform feature point matching on the feature points of the processed writing image and the feature points of the pre-shot image to obtain a matched feature point pair.

[0219] The generation module 802 is further configured to perform image fusion between the processed writing image and the pre-shot image based on the matched feature point pairs.

[0220] In some embodiments, the acquisition module 801 is further configured to perform feature point matching on the feature points of the processed writing image and the feature points of the pre-shot image to obtain a mapping matrix that maps the processed writing image to the pre-shot image.

[0221] The generation module 802 is further configured to perform image fusion between the processed writing image and the pre-shot image according to the mapping matrix.

[0222] In some embodiments, the acquisition module 801 is further configured to identify the target question region in the processed writing image, perform feature extraction processing on the target question region, and obtain the feature points of the target question region;

[0223] The acquisition module 801 is also used to extract features from the pre-shot question region in the pre-shot image to obtain feature points of the pre-shot question region.

[0224] In some embodiments, the acquisition module 801 is further configured to identify the target answer region in the processed writing image;

[0225] The acquisition module 801 is further configured to determine the mapping position of the target answer region in the pre-shot image based on the feature points of the target question region and the feature points of the pre-shot question region;

[0226] The generation module 802 is further configured to map the target answer region to the mapping position in the pre-shot image to obtain the target image.

[0227] In some embodiments, the generation module 802 is further configured to display the target image upon detecting a preset operation by the user.

[0228] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0229] It should be noted that, in the embodiments of this application... Figure 8 The module division shown in the illustrated writing content acquisition device is illustrative and represents only a logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or be integrated into one unit by two or more units. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.

[0230] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0231] This application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the methods described above.

[0232] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.

[0233] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.

[0234] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0235] In one embodiment, the written content acquisition device provided in this application can be implemented as a computer program, and the computer program can be implemented as follows: Figure 9The device operates on the computer device shown. The memory of the computer device can store the various program modules that make up the above-described apparatus. The computer program, composed of the various program modules, causes the processor to execute the steps of the methods in the various embodiments of this application described in this specification.

[0236] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0237] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.

[0238] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0239] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0240] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.

[0241] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0242] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.

[0243] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0244] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0245] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0246] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0247] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0248] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for obtaining written content, characterized in that, The method includes: Acquire pre-taken images, including pages of the target book; When the user is in a writing state, acquire a writing image including the book page written by the user, wherein the target book page includes the book page written by the user; Based on the pre-shot image and the writing image, a target image corresponding to the user's writing content is obtained.

2. The method according to claim 1, characterized in that, The acquisition of pre-captured images, including pages of the target book, includes: If a preview image containing any page of a book is captured, determine whether the page of the book in the preview image is obscured; If the book page in the preview image is not obstructed, the pre-shot image is obtained based on the preview image, and the target book page includes the book page in the preview image.

3. The method according to claim 1 or 2, characterized in that, The acquisition of pre-captured images, including pages of the target book, includes: If a preview image including any of the book pages is captured, determine whether the book corresponding to the book page in the preview image is in a stable state; When the book corresponding to the book page in the preview image is in the stable state, the pre-shot image is obtained based on the preview image, and the target book page includes the book page in the preview image.

4. The method according to claim 3, characterized in that, The step of obtaining the pre-shot image based on the preview image includes: Obtain the book page area from the preview image; The book page area is corrected to obtain the pre-shot image.

5. The method according to claim 4, characterized in that, The step of obtaining the book page area in the preview image includes: Identify multiple outline corner points of the book page in the preview image, and obtain the book page area based on the multiple outline corner points.

6. The method according to claim 4, characterized in that, The step of obtaining the book page area in the preview image includes: Detect the number of book pages in the preview image; If there are multiple book pages in the preview image, determine whether the multiple book pages belong to the same book; When the multiple book pages belong to the same book, obtain the multiple book page regions corresponding to the multiple book pages.

7. The method according to claim 6, characterized in that, The step of correcting the book page area to obtain the pre-shot image includes: The correction process is performed on each of the plurality of book page regions to obtain at least one pre-shot image corresponding to the plurality of book page regions.

8. The method according to claim 1, characterized in that, The step of acquiring a writing image including the book page written by the user while the user is in a writing state includes: If a preview image including the user's body parts is captured, detect whether the user is in a writing state; When the user is in the writing state, acquire a writing image including the book page written by the user.

9. The method according to claim 8, characterized in that, The step of obtaining the writing image, including the book page written by the user, includes: Obtain the book page area shown in the preview image where the user has written on the book page; The area of ​​the book page written by the user is corrected to obtain a corrected image, which is the writing image.

10. The method according to claim 9, characterized in that, After correcting the area of ​​the book page written by the user to obtain the corrected image, the method further includes: The writing image is cropped to obtain a processed writing image, which does not include the user's body parts; The processed writing image is fused with the pre-shot image to obtain the target image.

11. The method according to claim 10, characterized in that, The step of fusing the processed writing image with the pre-captured image to obtain the target image includes: Feature extraction is performed on the processed writing image and the pre-shot image respectively to obtain feature points of the processed writing image and feature points of the pre-shot image; Based on the feature points of the processed writing image and the feature points of the pre-shot image, the processed writing image and the pre-shot image are fused to obtain the target image.

12. The method according to claim 11, characterized in that, The step of fusing the processed writing image with the pre-captured image based on the feature points of the processed writing image and the feature points of the pre-captured image includes: Feature point matching is performed on the feature points of the processed writing image and the feature points of the pre-shot image to obtain matched feature point pairs; Based on the matched feature point pairs, the processed writing image and the pre-shot image are fused together.

13. The method according to claim 11, characterized in that, The step of fusing the processed writing image with the pre-captured image based on the feature points of the processed writing image and the feature points of the pre-captured image includes: Feature point matching is performed on the feature points of the processed writing image and the feature points of the pre-shot image to obtain a mapping matrix that maps the processed writing image to the pre-shot image; Based on the mapping matrix, the processed writing image and the pre-shot image are fused together.

14. The method according to claim 11, characterized in that, The pre-captured image also includes a pre-captured question area. The feature extraction process performed on the processed writing image and the pre-captured image respectively to obtain feature points of the processed writing image and the pre-captured image includes: Identify the target question region in the processed writing image, perform feature extraction processing on the target question region, and obtain the feature points of the target question region; Feature extraction is performed on the pre-shot question region in the pre-shot image to obtain the feature points of the pre-shot question region.

15. The method according to claim 14, characterized in that, The pre-captured image also includes a pre-captured answer region. The step of fusing the processed written image with the pre-captured image based on feature points of the processed written image and feature points of the pre-captured image to obtain the target image includes: Identify the target answer region in the processed writing image; Based on the feature points of the target question area and the feature points of the pre-shot question area, determine the mapping position of the target answer area in the pre-shot image; The target answer region is mapped to the mapping position in the pre-shot image to obtain the target image.

16. The method according to claim 11, characterized in that, After obtaining the target image corresponding to the user's writing content based on the pre-captured image and the writing image, the method further includes: If the user's preset action is detected, the target image is displayed.

17. A device for acquiring written content, characterized in that, include: The acquisition module is used to acquire pre-shot images, including pages of the target book; The acquisition module is also used to acquire a writing image including the book page written by the user when the user is in a writing state, wherein the target book page includes the book page written by the user; The generation module is used to obtain a target image corresponding to the user's writing content based on the pre-shot image and the writing image.

18. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 16.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 16.