Writing state detection method and device, electronic equipment and storage medium

By detecting the changes in the writing content of the target area in the multi-frame continuous image of the user in the learning scene, the problem of inaccurate and time-consuming judgment of students' writing status in the prior art is solved, and a more efficient and accurate judgment of the writing status is achieved.

CN120220169APending Publication Date: 2025-06-27GUANGDONG XIAOTIANCAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312226.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when judging whether a student is in a writing state, there are problems of inaccurate judgment and time-consuming.

Method used

By acquiring multiple frames of continuous images of the user in the learning scene, it is detected whether the writing content of the target area changes, which is the area where the user's pen tip is located. When the writing content changes, it is judged that the user is in the writing state.

Benefits of technology

It improves the accuracy of judging users' writing status, reduces time-consuming, and can record the user's writing time in real time, providing richer feedback on writing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220169A_ABST
    Figure CN120220169A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a writing state detection method and device, electronic equipment and a storage medium. The method comprises the steps that multiple frames of continuous images of a user in a learning scene are acquired; detecting whether the writing content of a target area in the multi-frame continuous image is changed or not, wherein the target area comprises the area where the pen point of the user is located; and under the condition that the written content changes, judging that the user is in a writing state. By implementing the embodiment of the invention, the writing state of the user can be judged by detecting the change of the writing content of the target area in the multi-frame continuous image, the accuracy of judging the writing state of the user can be improved, and the time consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of image processing, including but not limited to a writing state detection method, device, electronic device, and storage medium. Background Art

[0002] In the general trend of educational informatization, the accurate monitoring of students' learning status has become an important part of improving teaching quality and effect evaluation. Among them, judging whether a student is in a writing state is of great significance for understanding aspects such as the student's learning engagement, attention concentration, and learning efficiency.

[0003] In the related art, it is possible to collect data of a user in a learning scenario, such as images, videos, etc., for processing and analysis to identify and understand the information contained therein to judge whether the user is in a writing state. However, the solution based on recognition and understanding has problems of inaccurate judgment and time consumption. Summary of the Invention

[0004] In view of this, the writing state detection method, device, electronic device, and storage medium provided by the embodiments of the present application can improve the accuracy of judging the writing state of a user and reduce the time consumption.

[0005] The first aspect of the embodiments of the present application discloses a writing state detection method, which is applied to a terminal device. The method includes:

[0006] Obtain multiple consecutive frames of images of a user in a learning scenario;

[0007] Detect whether the writing content in the target area in the multiple consecutive frames of images changes, where the target area is the area including the tip of the user's pen;

[0008] In the case where the writing content changes, judge that the user is in a writing state.

[0009] In the above technical solution, the terminal device can judge the writing state of the user by detecting the change of the writing content in the target area in multiple consecutive frames of images, which can improve the accuracy of judging the writing state of the user and reduce the time consumption.

[0010] In some possible embodiments, the detecting whether the writing content in the target area in the multiple consecutive frames of images changes includes:

[0011] Identify the area image including the target area in each frame of the multiple consecutive frames of images;

[0012] Detect the feature similarity of the area images of adjacent frames of images in the multiple consecutive frames of images;

[0013] Wherein, if the feature similarity indicates that the features of the target region in adjacent frame images are not similar, the written content of the target region has changed.

[0014] In the above technical solution, by identifying the target region in multiple consecutive images and detecting the feature similarity of the target region in adjacent frame images, it is possible to determine whether the written content of the target region has changed, so as to efficiently and accurately detect whether the user is in a writing state.

[0015] In some possible embodiments, detecting the feature similarity of the regional images of adjacent frame images in the multiple consecutive images includes:

[0016] Inputting the regional images of the adjacent frame images into a preset similarity detection model to obtain the feature similarity of the adjacent frame images. The preset similarity detection model is obtained by training a first initial model based on sample image pairs and the feature similarities of the sample image pairs.

[0017] In the above technical solution, by using a preset similarity detection model to detect the feature similarity of adjacent frame images in multiple consecutive images, it is possible to quickly detect whether adjacent frame images or image features are similar, thereby further improving the accuracy and efficiency of writing state recognition and reducing the time consumption.

[0018] In some possible embodiments, each frame of the image includes the user's hand. Identifying the regional images including the target region in each frame of the multiple consecutive images includes:

[0019] Detecting the hand position corresponding to the hand in each frame of the image;

[0020] Obtaining the pen tip position according to the hand position;

[0021] Identifying the image within a preset range of the pen tip position as the regional image of the target region.

[0022] In the above technical solution, by detecting the hand position to assist in determining the pen tip position, the search range for finding the pen tip position and the target region in each frame of the image can be reduced, so that the target region can be more accurately identified, and the accuracy of target region recognition can be improved.

[0023] In some possible embodiments, detecting the hand position corresponding to the hand in each frame of the image includes:

[0024] Input each frame of the image into a preset target detection model to detect the hand position of the hand in each frame of the image. The target detection model is trained from a second initial model based on hand sample images and the corresponding hand positions of the hand sample images.

[0025] In the above technical solution, by detecting the hand position corresponding to the hand in each frame of the image through a preset target detection model, the search range for finding the pen tip position and the target area in each frame of the image can be reduced, the amount of calculation can be reduced, the interference of background information can be reduced, and the accuracy and efficiency of locating the target area in each frame of the image can be improved.

[0026] In some possible embodiments, obtaining the pen tip position of the pen tip according to the hand position includes:

[0027] Determine the area where the user's hand is located according to the hand position of the user;

[0028] Input the image including the area where the user's hand is located into a preset regression model to obtain the pen tip position of the pen tip in each frame of the image. The regression model is trained from a third initial model based on pen tip sample images and the corresponding pen tip positions of the pen tip sample images.

[0029] In the above technical solution, by inputting the image of the area where the hand is located into a preset regression model, the pen tip position of the pen tip in each frame of the image can be obtained quickly and accurately.

[0030] In some possible embodiments, the method further includes:

[0031] When the written content does not change, determine that the user is in a non-writing state.

[0032] In some possible embodiments, the method further includes:

[0033] When it is detected that the user is in the writing state, record the duration of the user being in the writing state.

[0034] In the above technical solution, by recording the writing duration of the user, a richer writing experience feedback can be provided for the user. According to the writing duration of the user, the terminal device can intelligently remind the user to rest or continue writing, enhancing the interactivity.

[0035] A second aspect of the embodiments of the present application discloses a writing state detection device, which is applied to a terminal device. The device includes:

[0036] An image acquisition module, configured to acquire multiple consecutive images of the user in a learning scenario;

[0037] A change detection module that detects whether the writing content in the target area of the multi-frame consecutive images changes, where the target area is the area including the position of the user's pen tip;

[0038] A status determination module that determines that the user is in a writing state when the writing content changes.

[0039] A third aspect of the embodiments of the present application discloses an electronic device, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor implements the method described above.

[0040] A fourth aspect of the embodiments of the present application discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented.

[0041] Compared with the related art, the embodiments of the present application at least include the following beneficial effects:

[0042] A writing state detection method, device, electronic device, and storage medium disclosed in the embodiments of the present application obtain multi-frame consecutive images of a user in a learning scenario, detect whether the writing content in the target area of the multi-frame consecutive images changes, where the target area is the area including the position of the user's pen tip, and determine that the user is in a writing state when the writing content changes. It can determine the user's writing state by detecting changes in the writing content in the multi-frame consecutive images, improve the accuracy of determining the user's writing state, and reduce the time consumption. Description of the Drawings

[0043] The accompanying drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.

[0044] Figure 1 It is a schematic flowchart of a writing state detection method in an embodiment;

[0045] Figure 2 It is a schematic flowchart of a target area content change detection method in an embodiment;

[0046] Figure 3 It is a schematic diagram of a target area image recognition method in an embodiment;

[0047] Figure 4 It is a schematic overall flowchart of a writing state detection method in an embodiment;

[0048] Figure 5 It is a block diagram of a writing state detection device in an embodiment;

[0049] Figure 6 It is a structural block diagram of an electronic device in an embodiment. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0052] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0053] It should be noted that the terms "first / second / third" involved in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0054] The embodiments of the present application disclose a writing state detection method, device, electronic device, and storage medium, which can improve the accuracy of judging the writing state of a user and reduce the time consumption. The following will be described in detail respectively.

[0055] It should be understood that the execution subject of each method embodiment of the present application is various types of terminal devices. For example, it can be a smart phone, a tablet computer, a wearable device, a learning machine, a tutoring machine, a learning tablet, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The specific type of the terminal device is not limited in the embodiments of the present application.

[0056] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a writing state detection method in an embodiment. This method is applied to a terminal device, such asFigure 1 As shown, the method may include the following steps:

[0057] Step 101, obtain multiple consecutive images of the user in the learning scenario.

[0058] In the embodiments of the present application, images of the surrounding environment where the terminal device is located can be captured by the camera of the terminal device. Exemplarily, the learning scenario may include various environments related to the user's learning behavior, such as a classroom scenario with desks, chairs, blackboards, projection devices, etc.; a home study scenario with desks, bookshelves, computers and other learning supplies; an online learning scenario where the user accesses an online course platform through an electronic device such as a computer, tablet, or mobile phone. The embodiments of the present application are not limited thereto.

[0059] As an alternative implementation, after the terminal device obtains the surrounding environment image, it can further confirm whether there is a user and locate the position of the user in the scenario through a face detection algorithm. Exemplarily, the face detection algorithm can scan each area in the image to identify the part with face features. At the same time, in order to determine whether the detected user is in the learning scenario, other information can also be combined for comprehensive judgment. For example, in the image area where a face is detected, whether there are learning-related items such as books and pens can be identified through an object detection algorithm.

[0060] In the case of determining that the user is in the learning scenario, the terminal device captures multiple consecutive images of the user in the learning scenario through the camera. In the embodiments of the present application, multiple consecutive images refer to a series of image frames that are closely connected in time and have a sequential order obtained by the terminal device. The image frames include at least two images, and each image frame includes relevant information of the user in the learning scenario. These image frames can be collected at regular time intervals, and the time interval between adjacent frames can be in seconds.

[0061] In some embodiments, the terminal device obtains multiple consecutive images of the user in the learning scenario, which can be triggered by the user's active operation. For example, the user can click on the learning monitoring application on the terminal device. With the user's consent, the terminal device starts the camera and obtains multiple consecutive images in the learning scenario. It can be understood that when the user does not perform an operation, the camera is in the closed state and will not capture images to protect the user's privacy.

[0062] In some other embodiments, the terminal device obtains multiple consecutive images of the user in the learning scenario, and can also be triggered by sensor assistance. For example, when the terminal device detects the presence of a user in the surrounding environment through an infrared sensor, the camera can be automatically started to capture images.

[0063] Step 102: Detect whether the writing content in the target area changes in multiple consecutive images. The target area is the area including the user's pen tip.

[0064] After the terminal device obtains multiple consecutive images of the user in the learning scenario, it can process and analyze the multiple consecutive images, locate the target area in the multiple consecutive images, and detect whether the writing content in the target area changes. The target area is the area including the user's pen tip. In some embodiments, it is first necessary to determine the target area, that is, the area where the pen tip is located. The target detection algorithm can be used to directly locate the pen tip in the multiple consecutive images, and then a fixed-size area centered on the pen tip is defined as the target area.

[0065] Since the pen tip may appear as a very small area in the image, its size is tiny and the features are not obvious. If the area where the pen tip is located is directly recognized from each frame of the image, it may be extremely vulnerable to factors such as image noise, complex background, and resolution, which may lead to great recognition difficulty and poor accuracy. And the writing behavior is usually completed by the user holding the pen with the hand, and the hand occupies a relatively large area in the image and has more recognizable features such as richer shapes and textures. Therefore, the area where the user's hand is located can be recognized first, and then the area where the user's pen tip is located can be recognized according to the regional image of the area where the hand is located.

[0066] Exemplarily, the terminal device can recognize the position of the user's hand in the multiple consecutive images through a preset target detection model, determine the area where the hand is located in the image according to the position of the user's hand, further recognize the position of the user's pen tip through a preset regression model after obtaining the image of the area where the hand is located, and then determine the target area centered on the pen tip position. Finally, compare whether the features of the regional images of the target area are similar. Optionally, a preset similarity detection model can be used to detect the feature similarity of the target area images of adjacent frame images, and determine whether the writing content in the target area changes according to the feature similarity.

[0067] Step 103: When the writing content changes, determine that the user is in the writing state.

[0068] In some embodiments, the change in the writing content means that the writing content increases or decreases. When the writing content increases, it is determined that the user is in the writing state; when the writing content decreases, it is considered that the user may have erased or modified the writing content, and it can also be determined that the user is in the writing state. In some other embodiments, the change in the writing content can also be defined by the user himself / herself, and the embodiments of the present application do not make specific limitations. Correspondingly, when the writing content does not change, it is determined that the user is in the non-writing state.

[0069] By adopting the above embodiments, it is possible to determine the writing state of the user by detecting the changes in the writing content of the target area in multiple consecutive images, which can improve the accuracy of judging the writing state of the user and reduce the time consumption.

[0070] In some embodiments, the writing state detection method further includes:

[0071] When it is detected that the user is in the writing state, record the duration of the user being in the writing state.

[0072] When the terminal device detects that the user is in the writing state, immediately obtain the current timestamp as the writing start time. During the period when the user is in the writing state, collect multiple consecutive images at a preset time interval, and determine whether the user is still in the writing state. When it is detected that the user changes from the writing state to the non-writing state, obtain the timestamp again as the writing end time. Then, subtract the start time stamp from the end time stamp to obtain the duration of the user being in the writing state.

[0073] By adopting the above embodiments, by detecting in real time whether the user is in the writing state and thus recording the writing duration of the user, it is possible to provide the user with a richer writing experience feedback. According to the writing duration of the user, the terminal device can intelligently remind the user to rest or continue writing, enhancing the interactivity.

[0074] In some embodiments, the multiple consecutive images include the topic that the user is currently writing. The writing state detection method further includes:

[0075] When it is determined that the user is in the writing state, determine the position of the topic that the user is currently writing according to the nib position of each frame of the image.

[0076] In the embodiments of the present application, when it is determined that the user is in the writing state, the OCR technology can be used to identify the text content on the book page in multiple consecutive images, divide the topic area, and assign a unique identifier and a bounding box to each topic area. According to the relative position of the nib position and the bounding box of the topic area, determine the topic area where the nib is located. If the nib position continuously appears within a certain topic area, it is considered that the user is writing this topic. In addition, the distance between the nib position and the center of the topic area can also be calculated, and the topic with the closest distance is the topic that the user is currently writing.

[0077] By adopting the above embodiments, when it is determined that the user is in the writing state, the position of the topic that the user is currently writing can also be located according to the nib position of each frame of the image, which can improve the user experience.

[0078] In the above step 102, to further illustrate how to detect whether the writing content in the target area of multiple consecutive images changes, please refer to Figure 2 ,Figure 2 It is a schematic flowchart of a method for detecting changes in the content of a target area in an embodiment. In some embodiments, the method may include the following steps:

[0079] Step 201, identify the area image including the target area in each frame of a multi-frame continuous image.

[0080] Exemplarily, the terminal device can process each frame of the image to identify the area image including the target area in each frame. The target area is the area where the user's pen tip is located. The area image is an extraction of a sub-region of each frame of the image, and its range and shape are based on the principle of being able to completely contain the target area. It can be a rectangle, a circle, or other irregular shapes. The embodiments of the present application do not make specific limitations here. It can be understood that the size of the area image including the target area in each frame of the image is the same.

[0081] In some embodiments, when the terminal device detects the presence of a pen in a multi-frame continuous image through a target detection algorithm, before identifying the area image including the target area in each frame of the image, it can first preprocess each frame of the image to improve the accuracy of subsequent identification. For example, use a Gaussian filtering algorithm or a median filtering algorithm, etc. to perform denoising processing on the image to reduce the influence of noise on the identification of the target area.

[0082] In the embodiments of the present application, the pen can be any type of writing tool, including a pencil, a ballpoint pen, a fountain pen, a writing brush, etc., or an electronic pen used on electronic devices such as a tablet computer or a graphics tablet. The embodiments of the present application do not make specific limitations. The area image of the area where the user's pen tip is located includes the user's writing content. By detecting whether the writing content has changed, it is possible to determine whether the user is in a writing state.

[0083] As described above, the target area in each frame of the image is the area where the pen tip is located. The area where the pen tip is located can be identified by first detecting the area of the hand in each frame of the image. In some embodiments, each frame of the image includes the user's hand. Identifying the area image including the target area in each frame of a multi-frame continuous image includes:

[0084] Detect the hand position corresponding to the hand in each frame of the image;

[0085] Obtain the pen tip position of the pen tip according to the hand position;

[0086] Identify the image within a preset range of the pen tip position as the area image of the target area.

[0087] The above method is exemplified by Figure 3 For example, please refer to Figure 3 , Figure 3Schematic diagram of the method for recognizing the regional image of the target area in an embodiment. In the embodiments of the present application, a plurality of consecutive images refer to a set of image frames arranged in sequence in a time series and having continuity in time and content between adjacent frames. The embodiments of the present application do not limit the number of frames of the collected images. In Figure 3 two collected images are taken as examples for illustration. As Figure 3 shown, in each frame of the image, first, the hand position corresponding to the user's hand in each frame of the image is detected to determine the position information of the user's hand in the image. For example, a rectangular frame can be used to mark the area where the hand position is located and display the area coordinates to represent the range of the hand in the image. Based on the position information of the hand in the image, the pen tip position is further determined. After determining the pen tip position, the area within the preset range of the pen tip position can be recognized as the target area.

[0088] In some embodiments, the detecting the hand position corresponding to the hand in each frame of the image includes:

[0089] Input each frame of the image into a preset target detection model to detect the hand position of the hand in each frame of the image. The target detection model is obtained by training a second initial model based on the hand sample image and the hand position corresponding to the hand sample image.

[0090] In the embodiments of the present application, the hand sample image is the image data used to train the second initial model and can include hand images in various scenarios. For example, static hand images taken from different angles, covering hand images from the front, side, and back, as well as images of the hand performing various actions, such as making a fist, opening, and holding an object. Moreover, in order to make the model have wide applicability, the hands in these sample images can come from people of different age groups, genders, and skin colors. By collecting a large number of hand sample images, the target detection model can learn various possible hand features.

[0091] The hand position corresponding to the hand sample image is the hand position annotation information matched in each hand sample image. An image annotation tool can be used to mark the area where the hand is located with a rectangular frame and mark the area coordinates on each hand sample image. For example, for a sample image of a hand writing, the marked rectangular frame should completely contain all the fingers and the palm part of the hand, and the boundary should be clear and accurate.

[0092] In some embodiments, the second initial model may be a YOLO (You Only Look Once) model, an SSD (Single Shot MultiBox Detector) model, a Faster R-CNN (Faster Region-based Convolutional Neural Network) model, an HRNet (High-Resolution Network) model, a RetinaNet (Retina Network) model, etc. The embodiments of the present application are not limited thereto.

[0093] During the training process, the second initial model extracts features from the input hand sample images to distinguish the feature patterns between the hand and other image contents, and adjusts the model parameters according to the difference between the predicted hand position and the actual labeled hand position. The trained second initial model is the target detection model. After inputting each frame of the image to be detected into the trained target detection model, the target detection model can output the position information of the hand in that frame of the image. For example, it can be presented in the form of a rectangular box, that is, a rectangular box showing the area where the labeled hand position is located, as well as the upper left coordinate (x1, y1) and the lower right coordinate (x2, y2) of the rectangular box in the image. Through these two sets of coordinates, the hand position corresponding to the hand in each frame of the image can be determined.

[0094] By adopting the above embodiments, detecting the hand position corresponding to the hand in each frame of the image through the preset target detection model can narrow the search range for finding the pen tip position and the target area in each frame of the image, reduce the calculation amount and the interference of background information, and improve the accuracy and efficiency of locating the target area in each frame of the image.

[0095] After the terminal device detects the hand position corresponding to the hand in each frame of the image, it can obtain the pen tip position of the pen according to the hand position. It can be understood that when the user holds the pen for writing, there is a relatively fixed spatial position relationship between the pen and the hand. If the pen tip position cannot be obtained according to the hand position, it means that the user is not holding the pen, so it can be determined that the user is in a non-writing state. If the pen tip position can be obtained according to the hand position, the target area in each frame of the image can be further recognized.

[0096] In some embodiments, the obtaining the pen tip position of the pen according to the hand position includes:

[0097] Determining the area where the user's hand is located according to the user's hand position;

[0098] Input an image including the area where the user's hand is located into a preset regression model to obtain the nib position of the nib for each frame of the image. The regression model is obtained by training a third initial model based on nib sample images and the corresponding nib positions of the nib sample images.

[0099] In the embodiments of the present application, after detecting the hand position corresponding to the user's hand in each frame of the image through the target detection model, the area where the hand is located can be marked using a rectangular box, that is, the area where the user's hand is located is determined. The terminal device can extract the area image of the area where the user's hand is located from each frame of the image. The area image of the area where the user's hand is located is the image within the rectangular area defined by the hand position bounding box output by the target detection model. This image includes the hand and may also include some background parts around the hand.

[0100] The preset regression model is used to predict the nib position of the nib in the input image information. The regression model is a machine learning model, and its core goal is to establish a quantitative relationship between independent variables and dependent variables to predict continuous numerical values. In the scenario of predicting the nib position in an image, the independent variable is the nib feature extracted from the image, and the dependent variable is the nib position of the nib. Exemplarily, the nib position can be represented in the form of two-dimensional coordinates. The regression model can accurately output the nib position of the nib in the image to be detected by learning the mapping relationship between the nib features and the nib positions in the nib sample images.

[0101] In some embodiments, the nib sample image refers to a series of images collected that contain nibs. These images can be captured under various conditions such as different angles and different lighting conditions, with the aim of covering images of the nib in various possible states. It can be understood that the nibs in the nib sample images can include fountain pen nibs, ballpoint pen nibs, pencil nibs, brush nibs, stylus nibs, etc., which are not limited in the embodiments of the present application.

[0102] The nib position corresponding to the nib sample image refers to the specific coordinate position information of the nib marked in each nib sample image in that image. Usually in pixels, the position of the nib in the image is represented by two-dimensional coordinates (x, y), where x represents the position in the horizontal direction and y represents the position in the vertical direction.

[0103] In some embodiments, the third initial model may be a linear regression model, a polynomial regression model, a support vector regression model, a random forest regression model, a neural network regression model, etc., and the embodiments of the present application do not make any limitations. During the training process of the third initial model, the third initial model will make a prediction based on the input pen tip sample image, output an estimated value of the pen tip position, and then compare the estimated value with the pen tip position corresponding to the pen tip sample image. The parameters of the third initial model are adjusted by calculating the difference value between the two. The trained third initial model is the regression model. The regional image of the area where the hand is located in each frame of the image is input into the regression model, and the regression model can output the pen tip position of the pen tip in each frame of the image. Exemplarily, the pen tip position can be marked in the form of two-dimensional coordinates.

[0104] By adopting the above embodiments, by inputting the image of the area where the hand is located into the preset regression model, the pen tip position of the pen tip in each frame of the image can be quickly and accurately obtained.

[0105] In the embodiments of the present application, when the terminal device obtains the pen tip position, it further identifies the image within the preset range of the pen tip position as the regional image of the target area.

[0106] The preset range is a pre-set parameter. In some embodiments, the preset range can be a fixed pixel distance. For example, centered on the pen tip position, the range is 50 pixels around; it can also be a proportional value, such as centered on the pen tip position, the range is 10% of the image width and height.

[0107] In some embodiments, the preset range can have different shapes and setting methods. For example, centered on the pen tip position, a fixed pixel value is set as the radius to form a circular area to obtain the target area, or centered on the pen tip position, a fixed pixel value of length and width is set to form a rectangular area to obtain the target area.

[0108] By setting the range value, the terminal device can intercept the image within the preset range from the regional image of the area where the hand is located, and the image within the preset range is the regional image of the target area.

[0109] After the terminal device completes the above-mentioned hand position detection, pen tip position acquisition, and target area image recognition for the first frame of the image, for the second frame or more frames of the image, the same operation process is repeated. After the terminal device obtains the regional image of the target area in each frame of the image, it executes the method described in step 202.

[0110] Step 202, detect the feature similarity of the regional images of adjacent frames in multiple consecutive frames of images;

[0111] Among them, if the feature similarity indicates that the features of the target region in adjacent frame images are not similar, the written content of the target region changes.

[0112] In a series of consecutive images, except for the first image and the last image, each image has a previous adjacent image and a next adjacent image. Adjacent frame images refer to two images that are closely connected in chronological order. In some possible embodiments, detecting the feature similarity of the region images of adjacent frame images in a series of consecutive images may include at least the following two exemplary methods.

[0113] The first example:

[0114] In some embodiments, the detecting the feature similarity of the region images of adjacent frame images in a series of consecutive images includes:

[0115] Input the region images of adjacent frame images into a preset similarity detection model to obtain the feature similarity of the adjacent frame images. The preset similarity detection model is obtained by training a first initial model based on sample image pairs and the feature similarities of the sample image pairs.

[0116] In the embodiments of the present application, the similarity detection model is used to detect the feature similarity of adjacent frame images. Optionally, the feature similarity is a numerical index used to indicate whether two images are similar at the feature level. The value of the feature similarity can be 0 or 1. When the value of the feature similarity output by the similarity detection model is 0, it means that the adjacent two frame images are not similar. It should be noted that as long as the similarity detection model detects a difference in the features of the adjacent two frame images, it means that the input adjacent two frame images are not similar; when the value of the feature similarity output by the similarity detection model is 1, it means that the adjacent two frame images are similar, that is, the similarity detection model detects that there is no difference in the features of the adjacent two frame images.

[0117] A sample image pair refers to the image data used to train the first initial model. That is, the input of the first initial model is a pair of sample images, and the output is the feature similarity predicted by the model between the sample image pair. Each pair of sample images can be similar or dissimilar.

[0118] The feature similarity of a sample image pair refers to the result of annotating or calculating the similarity of the sample image pair. Similarly, the feature similarity of a sample image pair can take a value of 0 or 1. When the value of the feature similarity of the sample image pair is 0, it means that the sample image pair is not similar; when the value of the feature similarity of the sample image pair is 1, it means that the sample image pair is similar.

[0119] Optionally, the first initial model may be a convolutional neural network (CNN), a Siamese Network, a Transformer architecture, etc., which is not limited in the embodiments of the present application.

[0120] In some embodiments, when the sample image pair is input into the first initial model, the model processes the input image pair according to its current parameters and outputs a predicted feature similarity. Then, the predicted value is compared with the actual feature similarity (i.e., the label) of the sample image pair, the error between the two is calculated, and the parameters of the model are adjusted according to the error, so that the feature similarity predicted by the first initial model gradually approaches the true feature similarity. The trained first initial model is the similarity detection model.

[0121] Second example:

[0122] In some embodiments, detecting the feature similarity of the regional images of adjacent frames in multiple consecutive images includes:

[0123] Obtain the features of the regional images of adjacent frames, and obtain the feature similarity of the regional images of adjacent frames according to the features of the regional images of adjacent frames.

[0124] Exemplarily, a feature extraction algorithm may be used to extract the features of the regional images of each frame. For example, the Scale-Invariant Feature Transform (SIFT) algorithm, the Speeded-Up Robust Features (SURF) algorithm, the edge detection algorithm, etc. may be used, and the extracted features may be color features, texture features, shape features, etc.

[0125] After the features of the regional images of each frame are extracted, the feature similarity of the regional images of adjacent frames may be calculated. For example, the Euclidean distance between two feature vectors may be calculated. The smaller the distance, the higher the similarity. Assume that the two feature vectors are X = (x1, x2,... x n ) and Y = (y1, y2,... y n ), and their Euclidean distance may be expressed as follows:

[0126]

[0127] where n represents the dimension of the feature vector, and i represents the index variable used to traverse each element in the feature vector. When the Euclidean distance value between the feature vectors of adjacent frames is 0, it indicates that the features of the regional images of adjacent frames are similar; when the Euclidean distance value between the feature vectors of adjacent frames is not 0, it indicates that the features of the regional images of adjacent frames are not similar.

[0128] In some embodiments, the feature similarity can also be measured by calculating the cosine similarity between two feature vectors. The closer the cosine value between two feature vectors is to 1, the more similar the two vectors are. The calculation formula can be expressed as follows:

[0129]

[0130] Where X·Y represents the dot product of two feature vectors, and ||X|| and ||Y|| represent the norms of the two feature vectors respectively. When the cosine similarity between the feature vectors of adjacent frame images is 1, it indicates that the features of the regional images of the adjacent frame images are similar; when the cosine similarity between the feature vectors of adjacent frame images is not 1, it indicates that the features of the regional images of the adjacent frame images are not similar.

[0131] By adopting the above embodiments, detecting the feature similarity of adjacent frame images in multiple consecutive frames of images through a preset similarity detection model can quickly detect whether the adjacent frame images or image features are similar, thereby further improving the accuracy and efficiency of writing state recognition and reducing the time consumption.

[0132] After obtaining the feature similarity of the regional images of adjacent frame images, if the feature similarity indicates that the features of the target region of the adjacent frame images are not similar, then the writing content of the target region has changed.

[0133] In some embodiments, if the multiple consecutive frames of images collected only include two frames of images, then when it is detected that the features of the regional images of the target region in the two frames of images are not similar, it indicates that the writing content of the target region has changed; if the multiple consecutive frames of images collected include three frames of images or more frames of images, then as long as it is detected that the features of the regional images of the target region in a group of images are not similar, it indicates that the writing content of the target region has changed, and a group refers to two adjacent frames of images as a group.

[0134] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the overall process of the writing state detection method in an embodiment. As Figure 4 shown, the method may include the following steps:

[0135] Step 401, obtain multiple consecutive frames of images of the user in the learning scenario.

[0136] For the obtaining of multiple consecutive frames of images of the user in the learning scenario described in step 401, the specific implementation manner may refer to the relevant examples in the above step 101, and will not be elaborated here.

[0137] Step 402, detect the hand position corresponding to the hand in each frame of image.

[0138] In step 402, the hand position corresponding to the hand in each frame of the image is detected by inputting each frame of the image into a preset target detection model to obtain the hand position of the hand in each frame of the image. For the specific implementation, reference can be made to the above related embodiments and will not be elaborated here.

[0139] Step 403: Detect the pen tip position of the pen tip in each frame of the image.

[0140] When the hand position corresponding to the hand in each frame of the image is obtained, the terminal device can further obtain the pen tip position of the pen tip in each frame of the image according to the hand position. By inputting the image of the area where the user's hand is located into a preset regression model, the pen tip position of the pen tip in each frame of the image can be obtained. For the specific implementation, reference can be made to the above related embodiments and will not be elaborated here.

[0141] Step 404: Identify the area image of the target area as the image within the preset range of the pen tip position.

[0142] When the pen tip position of the pen tip in each frame of the image is detected, with the pen tip position as the center, the image within the preset range of the pen tip position is the area image of the target area. For the specific implementation, reference can be made to the above related embodiments and will not be elaborated here.

[0143] Step 405: Determine whether the user is in a writing state according to whether the writing content in the target area has changed.

[0144] In some embodiments, to determine whether the writing content in the target area has changed, it can be achieved by calculating the feature similarity of the area images of adjacent frames in multiple consecutive frames of images. For the specific implementation, reference can be made to the above related embodiments and will not be elaborated here. When the writing content in the target area has changed, it is determined that the user is in a writing state; when the writing content in the target area has not changed, it is determined that the user is in a non-writing state.

[0145] By adopting the above embodiments, it is possible to determine the user's writing state by detecting the change of the writing content in the target area in multiple consecutive images, which can improve the accuracy of judging the user's writing state and reduce the time consumption.

[0146] Please refer to Figure 5 , Figure 5 which is a block diagram of a writing state detection device in an embodiment. As Figure 5 shown, the writing state detection device 500 includes: an image acquisition module 510, a change detection module 520, and a state judgment module 530, where:

[0147] The image acquisition module 510 is configured to acquire multiple consecutive frames of images of the user in the learning scenario;

[0148] A change detection module 520 detects whether the writing content in the target area in the multi-frame consecutive images changes, where the target area is the area including the position of the user's pen tip;

[0149] A state judgment module 530 is used to judge that the user is in a writing state when the writing content changes.

[0150] In some embodiments, the change detection module 520 is specifically configured to:

[0151] Identify the regional image including the target area in each frame of the multi-frame consecutive images;

[0152] Detect the feature similarity of the regional images of adjacent frame images in the multi-frame consecutive images;

[0153] Wherein, if the feature similarity indicates that the features of the target area in adjacent frame images are not similar, the writing content in the target area changes.

[0154] In some embodiments, the change detection module 520 is specifically configured to:

[0155] Input the regional images of the adjacent frame images into a preset similarity detection model to obtain the feature similarity of the adjacent frame images, and the preset similarity detection model is obtained by training a first initial model according to a sample image pair and the feature similarity of the sample image pair.

[0156] In some embodiments, each frame of image includes the user's hand, and the change detection module 520 is specifically configured to:

[0157] Detect the hand position corresponding to the hand in each frame of image;

[0158] Obtain the pen tip position of the pen tip according to the hand position;

[0159] Identify the image within a preset range of the pen tip position as the regional image of the target area.

[0160] In some embodiments, the change detection module 520 is further configured to:

[0161] Input each frame of image into a preset target detection model to detect the hand position of the hand in each frame of image, and the target detection model is obtained by training a second initial model according to a hand sample image and the hand position corresponding to the hand sample image.

[0162] In some embodiments, the change detection module 520 is further configured to:

[0163] Determine the area where the user's hand is located according to the position of the user's hand;

[0164] Input an image including the area where the user's hand is located into a preset regression model to obtain the nib position of the nib for each frame of the image. The regression model is obtained by training a third initial model based on nib sample images and the corresponding nib positions of the nib sample images.

[0165] In some embodiments, the state determination module 530 is further configured to:

[0166] When the written content does not change, determine that the user is in a non-writing state.

[0167] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0168] It should be noted that in the embodiments of the present application Figure 6 The division of modules of the writing state detection device shown is schematic, merely a logical function division. In actual implementation, there may be other division methods. In addition, each functional unit in the various embodiments of the present application may be integrated in one processing unit, may exist separately physically, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware, or in the form of software functional units, or in the form of a combination of software and hardware.

[0169] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0170] Please refer to Figure 6 , Figure 6 which is a structural block diagram of an electronic device in an embodiment. As Figure 6As shown, the electronic device 600 may include: a processor 610, a memory 620, and a bus 630.

[0171] Among them, the processor 610 calls the executable program code stored in the memory 620 and executes any one of the writing state detection methods disclosed in the embodiments of the present application. Those skilled in the art can understand that Figure 6 the structure of the electronic device shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0172] The processor 610 is used to execute to obtain multiple consecutive frames of images of the user in the learning scenario; detect whether the writing content in the target area of the multiple consecutive frames of images changes, and the target area is the area including the tip of the user's pen; in the case where the writing content changes, determine that the user is in the writing state. For specific details, please refer to the detailed description in the method example, which will not be elaborated here.

[0173] In the embodiments of the present application, the processor 610 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor, or executed by a combination of hardware and software modules in the processor.

[0174] The memory 620 can be used to store software programs and modules. The processor 610 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 620. The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0175] In the embodiments of the present application, the processor 610 and the memory 620 are connected through the bus 630. The bus 630 is Figure 6 shown by a thick line in. The connection methods between other components are only for illustrative purposes and are not to be taken as a limitation. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.

[0176] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.

[0177] An embodiment of the present application provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the steps in the method provided in the above method embodiment.

[0178] Those of ordinary skill in the art can understand that all or part of the processes in implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a ROM, etc.

[0179] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium, storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0180] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" or "in some embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution is prior or subsequent. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The above descriptions of each embodiment tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0181] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, object A and / or object B can represent: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0182] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0183] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0184] The modules described above as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules; they can be located in one place or distributed to multiple network units; some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0185] In addition, in each embodiment of this application, each functional module can be fully integrated in a processing unit, or each module can be separately used as a unit, or two or more modules can be integrated in a unit; the above-mentioned integrated modules can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0186] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage media include: various media such as removable storage devices, read-only memory (ROM), magnetic disks or optical discs that can store program codes.

[0187] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as removable storage devices, ROMs, magnetic disks, or optical discs.

[0188] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0189] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0190] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0191] The above is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A writing state detection method, characterized in that: Applied to a terminal device, the method comprises: Acquire multiple frames of continuous images of the user in the learning scene; Detecting whether the written content of a target area in the plurality of continuous frames of images changes, the target area being an area including the pen tip of the user; When the written content changes, it is determined that the user is in a writing state.

2. The method according to claim 1, characterized in that The detecting whether the written content of the target area in the plurality of continuous frames of images changes comprises: Identifying a region image including the target region in each frame of the plurality of continuous frames of images; Detecting feature similarity of the regional images of adjacent frame images in the plurality of continuous frames of images; If the feature similarity indicates that the features of the target area of ​​adjacent frame images are not similar, the written content of the target area changes.

3. The method according to claim 2, characterized in that The detecting the feature similarity of the regional images of adjacent frame images in the plurality of continuous frames of images comprises: The regional images of the adjacent frame images are input into a preset similarity detection model to obtain the feature similarity of the adjacent frame images. The preset similarity detection model is obtained by training a first initial model based on a sample image pair and the feature similarity of the sample image pair.

4. The method according to claim 2, characterized in that: Each frame of the image includes the user's hand, and the identifying the region image including the target region in each frame of the plurality of continuous images includes: Detecting a hand position corresponding to the hand in each frame of the image; Acquiring a pen tip position of the pen tip according to the hand position; An image located within a preset range of the pen tip position is identified as a regional image of the target area.

5. The method according to claim 4, characterized in that The detecting the hand position corresponding to the hand in each frame of the image includes: Each frame of image is input into a preset target detection model to detect the hand position of the hand in each frame of image. The target detection model is obtained by training a second initial model based on hand sample images and the hand positions corresponding to the hand sample images.

6. The method according to claim 4, characterized in that The obtaining the pen tip position of the pen tip according to the hand position includes: Determining the area where the user's hand is located according to the position of the user's hand; An image of the area where the user's hand is located is input into a preset regression model to obtain the pen tip position of each frame image. The regression model is obtained by training a third initial model based on the pen tip sample image and the pen tip position corresponding to the pen tip sample image.

7. The method according to claim 1, characterized in that The method further comprises: When the written content does not change, it is determined that the user is in a non-writing state.

8. The method according to claim 1, characterized in that: The method further comprises: When it is detected that the user is in the writing state, the duration for which the user is in the writing state is recorded.

9. A writing state detection device, characterized in that: Applied to a terminal device, the device comprises: An image acquisition module is used to acquire multiple frames of continuous images of the user in the learning scene; a change detection module, detecting whether the written content of a target area in the plurality of consecutive frames of images has changed, the target area being an area including the pen tip of the user; The state judgment module is used to judge whether the user is in a writing state when the writing content changes.

10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor implements the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.