Image processing method and device, equipment, storage medium and program product
By acquiring and processing handwritten text images of the subjects, removing background information and performing segmentation and feature extraction, the accuracy and comprehensiveness issues of ADHD screening in existing technologies are solved, and low-cost and efficient ADHD detection is achieved.
Patent Information
- Application Number
- CN202410330352.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies for screening attention deficit hyperactivity disorder (ADHD) rely on external factors, resulting in high costs, high dependence, low accuracy and comprehensiveness, and require the support of wearable devices and controlled environments.
By obtaining the subject's handwritten continuous text image, removing background information, performing segmentation and feature extraction, and using single-word image features to determine ADHD, the dependence on the environment and equipment is reduced.
It improves the accuracy and comprehensiveness of ADHD screening, reduces costs, reduces dependence on external factors, and simplifies the testing process.
Smart Images

Figure CN120689939A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image processing method, apparatus, device, storage medium, and program product. Background Art
[0002] In related technologies, determining whether a subject suffers from Attention Deficit Hyperactivity Disorder (ADHD) often requires the assistance of wearable devices, the subject's cooperation, the test environment, resource support, and behavioral indicator assessments of the subject. Relying on external factors to determine whether a subject suffers from ADHD results in high costs, high dependence, and low applicability for screening whether the subject suffers from ADHD, making it impossible to guarantee the accuracy and comprehensiveness of whether the subject suffers from ADHD. Summary of the Invention
[0003] The present application provides an image processing method, apparatus, device, storage medium and program product.
[0004] The technical solution of this application is achieved as follows:
[0005] In a first aspect, the present application provides an image processing method, comprising:
[0006] Acquire a first text image containing continuous characters written by the subject;
[0007] removing background information from the first text image to obtain a second text image; the background information includes: a grid or horizontal lines;
[0008] Segmenting the continuous characters in the second text image to obtain at least two target single-character images;
[0009] Feature extraction is performed on each target single-word image to obtain single-word image features; the single-word image features are used to determine whether the subject has attention deficit hyperactivity disorder (ADHD).
[0010] In a second aspect, the present application provides an image processing device, comprising: an acquisition module, a removal module, a segmentation module, and an extraction module; wherein:
[0011] An acquisition module, configured to acquire a first text image containing continuous characters handwritten by the subject;
[0012] a removal module, configured to remove background information from the first text image to obtain a second text image; the background information includes: a grid or horizontal lines;
[0013] a segmentation module, configured to segment the continuous characters in the second text image to obtain at least two target single-character images;
[0014] The extraction module is used to extract features from each target single-word image to obtain single-word image features; the single-word image features are used to determine whether the subject has attention deficit hyperactivity disorder (ADHD).
[0015] In a third aspect, the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor to execute the image processing method provided above.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program enables a computer to execute the image processing method provided above.
[0017] In a fifth aspect, the present application provides a computer program product, comprising computer program instructions, which, when executed by a processor, implement the image processing method described in the first aspect.
[0018] In the technical solution provided in the present application, a first text image containing continuous handwritten text of the subject and the information of the subject are obtained; the grid or horizontal lines in the first text image are removed to obtain a second text image; the second text image is segmented to obtain a target single-word image; then the single-word image features in the target single-word image are extracted; and it is determined whether the subject suffers from ADHD based on the single-word image features. In this way, on the one hand, only the subject's handwritten continuous text image is needed to determine whether the subject suffers from ADHD, eliminating the dependence on external factors such as the environment, wearable devices, and the subject's cooperation when detecting whether the subject suffers from ADHD, saving costs and improving applicability; on the other hand, by segmenting the continuous handwritten text image, single-word images and single-word image features are obtained, and each single-word image feature is used to characterize the written text features of the subject, thereby determining whether the subject suffers from ADHD, further improving the accuracy and comprehensiveness of determining whether the subject suffers from ADHD. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 An image processing method provided in this embodiment of the application Figure 1 ;
[0020] Figure 2 An image processing method provided in this embodiment of the application Figure 2 ;
[0021] Figure 3 An image processing method provided in this embodiment of the application Figure 3 ;
[0022] Figure 4 An image processing method provided in this embodiment of the application Figure 4 ;
[0023] Figure 5 An image processing method provided in this embodiment of the application Figure 5 ;
[0024] Figure 6 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;
[0025] Figure 7 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] To facilitate understanding of the technical solutions of the embodiments of the present application, the relevant technologies of the embodiments of the present application are described below. The following relevant technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0028] In addition, in the embodiments of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used for a specific order or sequence.
[0029] The image processing method provided in the embodiments of the present application can be applied to the screening of Attention Deficit Hyperactivity Disorder (ADHD) in the medical field.
[0030] ADHD is a neurodevelopmental disorder that begins in childhood and is characterized by developmentally inappropriate attention deficits, hyperactivity, and impulsivity. Symptoms persist into adolescence in 60% to 80% of children, and affect adulthood in 50%. Clinical manifestations include significant difficulty concentrating, a short attention span, hyperactivity, impulsivity, irritability, and learning difficulties compared to children of the same age.
[0031] The national prevalence of ADHD among children is 5.6%, and it's trending upwards annually. If not promptly identified and intervened, it can have long-term negative consequences for children's learning, social, and emotional development. However, ADHD symptoms can be subtle and difficult to identify. Early screening is crucial to identify children with ADHD as early as possible, allowing for timely intervention and mitigating its impact on children.
[0032] Currently, clinicians assess ADHD through subjective descriptions of patients' behaviors. Assessment methods include medical clinical observation, examination-style interviews, and clinical scales. The limitations of these methods are that patients may have problems such as misunderstanding of questions and state recall bias. These methods have certain requirements for the cognitive level of the test subjects, are highly subjective, lack objectivity and consistency, are time-consuming, and have complex scoring methods. They also require a high level of professionalism from the diagnostician.
[0033] In recent years, the exploration of objective quantitative screening technologies for ADHD has continued to develop, including sustained attention performance tests, electroencephalogram (EEG), eye tracking, and body movement monitoring. These technologies provide more accurate and reliable methods for assessing ADHD through objective biometric data.
[0034] The sustained attention performance test usually takes 25 minutes, during which professional doctors test the child's visual and auditory integration function, attention, reaction control ability, etc. It is often used as an auxiliary diagnosis for ADHD.
[0035] EEG technology is a commonly used electrophysiological measurement method for the brain, capturing electrical activity through electrodes placed on the scalp. Specific implementation methods include recording children's EEG signals during specific tasks (such as continuous performance tasks) and using methods such as spectral analysis and spatiotemporal correlation to extract features for diagnostic assessment.
[0036] Eye tracking technology can be used to measure children's eye movements during visual attention tasks. Children with ADHD exhibit different eye movement patterns than typically developing children in target selection, response inhibition, and sustained attention. In practice, eye tracking or cameras are used to record children's eye movements during visual stimulation tasks, such as gaze duration and gaze shifts, to infer the likelihood of ADHD.
[0037] Activity monitoring technology uses wearable motion sensors or accelerometers to record children's movement data during daily activities or specific tasks, which is used to assess their movement activity level and restlessness.
[0038] However, the existing related technologies have the following technical problems:
[0039] (1) Low operability and acceptance of active monitoring: Sustained attention performance tests, EEG, and eye tracking technologies are task-driven active monitoring. Completing the test tasks requires time and effort, and requires the child's full participation and cooperation. This will affect the subject's compliance, the completion of the test process, and the acceptance of the screening work.
[0040] (2) Low universality and applicability of wearable devices: EEG and body movement monitoring technologies require subjects to wear contact devices, which may cause some discomfort and pressure to the children. In addition, families are usually not equipped with such wearable devices, and the cost of equipping such devices in screening work is high.
[0041] (3) Controlled testing environment affects the sensitivity of screening: Sustained attention performance tests, EEG, and eye tracking technology need to be conducted in a controlled laboratory environment. This specific environment may not fully represent the behavior of children in their daily lives. It lacks sensitivity to children's daily behavior.
[0042] (4) The need for medical resources will lead to limited screening popularity: During the continuous attention performance test and EEG test, professional doctors are required to provide guidance throughout the process, which will consume a certain amount of medical human resources. Occasionally, medical personnel will be occupied, resulting in the inability to conduct immediate testing, which will limit the popularity of screening.
[0043] (5) Complementary technologies are needed to improve the accuracy and comprehensiveness of screening: Existing physiological indicators (such as EEG, eye movements, and body movements) lack a way to assess patients' daily behavior and functionality.
[0044] In summary, determining whether a subject has ADHD often requires the assistance of external factors, such as wearable devices, the subject's cooperation, the testing environment, resource support, and behavioral assessments. This reliance on external factors results in high costs, high dependency, and low applicability, making it impossible to accurately and comprehensively determine whether a subject has ADHD.
[0045] Based on the above-mentioned related issues, an embodiment of the present application provides an image processing method. On the one hand, only a continuous handwritten text image of the subject is required to determine whether the subject suffers from ADHD, eliminating the dependence of the detection of whether the subject suffers from ADHD on external factors such as the environment, wearable devices, and the cooperation of the subject, saving costs and improving applicability; on the other hand, by segmenting the continuous handwritten text image, single-word images and single-word image features are obtained, and each single-word image feature is used to characterize the written text features of the subject, thereby determining whether the subject suffers from ADHD, further improving the accuracy and comprehensiveness of determining whether the subject suffers from ADHD.
[0046] The image processing method provided in the embodiments of the present application can be applied to fields including but not limited to image processing, medicine, etc.
[0047] The image processing method provided in the embodiments of the present application is executed by an image processing device and an electronic device, wherein the image processing device can be stored in the electronic device in the form of a software function model, the image processing device can also be integrated into the electronic device as a hardware function module, and the image processing device can also be combined with the electronic device in software and hardware to implement the image processing method. This application does not impose any restrictions on this.
[0048] In an embodiment of the present application, the electronic device may be a server, which may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The embodiment of the present application does not limit this.
[0049] In addition, the electronic device can also be a terminal device, which can be a mobile phone, a tablet personal computer (TPC), a media player, a smart TV, a laptop computer (LC), a personal digital assistant (PDA), a personal computer (PC), a camera, a video camera, a smart watch, a wearable device (WD) or an autonomous driving vehicle, etc., and the embodiments of the present application are not limited to this.
[0050] The present application provides an image processing method, such as Figure 1 As shown, the method includes the following steps S110 to S150:
[0051] Step S110: Acquire a first text image containing continuous characters handwritten by the subject.
[0052] In the embodiment of the present application, the subject may be a subject whose diagnosis or diagnosis is to be determined as to whether the subject suffers from ADHD.
[0053] The first text image can be a directly acquired original image or a pre-processed image. Pre-processing can include enhancement and correction. Generally, enhancement methods include, but are not limited to, histogram equalization, filtering, and image enhancement; correction methods include, but are not limited to, Hough line detection and Radon angle transformation.
[0054] In an embodiment of the present application, by using some algorithmic tools to enhance and correct the original text image, on the one hand, problems such as uneven lighting and noise in the original text image can be eliminated, making the handwritten text in the image clearer; on the other hand, the perspective tilt existing when taking the photo can be corrected through correction processing, so that the continuous text area becomes a horizontal and vertical rectangle, and the corrected rectangular composition area is cut out, making it easier to conduct subsequent evaluation of the neatness of the writing, thereby improving the accuracy of determining the object type.
[0055] The first text image can be an image of a handwritten composition or an image of an excerpted article. The first text image can be an image captured by an electronic device using a capture device, such as a scanned paper document or a handwritten composition captured using a camera. It can also be a video frame captured from a video file, or an image containing continuous text downloaded from a website. The first text image should be high-resolution and clear to preserve details and provide a basis for accurate analysis results.
[0056] Step S120: removing background information from the first text image to obtain a second text image; the background information includes: a grid or horizontal lines.
[0057] It should be understood that different formats of paper include different backgrounds. For example, the background of grid paper is a plurality of grids, and the background of lined paper is a plurality of lined lines.
[0058] In the embodiment of the present application, the removal of background information can be achieved by using morphological operations, connected domain detection, watershed algorithm, deep learning and other methods.
[0059] Step S130: Segment the continuous characters in the second text image to obtain at least two target single-character images.
[0060] It should be understood that the target single-word image contains only one word.
[0061] Here, the methods used for segmentation processing may include but are not limited to: segmentation based on connected regions, converting the handwritten Chinese character image into a binary image, and then using connected region analysis technology to divide the single word part of each Chinese character into different connected regions, and further screening and judgment can be performed according to the properties of the connected regions (such as area, width, height, etc.); segmentation based on projection, converting the handwritten Chinese character image into a binary image, and calculating the projection in the horizontal and vertical directions, and by detecting the peaks and valleys on the projection, the boundary position of the Chinese character can be determined and segmented; segmentation based on machine learning or deep learning, using machine learning (or deep learning) methods to train the model on a large number of annotated handwritten Chinese character data sets, thereby realizing automatic single-word segmentation.
[0062] In another embodiment of the present application, in addition to extracting single-word image features, global feature extraction can also be performed directly on the first text image, and the global feature extraction can be used to further quantitatively analyze and evaluate the neatness of the handwritten text. The method of extracting the global features of the first text image may include, but is not limited to: Histogram of Oriented Gradients (HOG), Principal Component Analysis (PCA), and wavelet transform features.
[0063] HOG features capture local shape information in an image. For handwriting, HOG features can describe the direction and strength of strokes, two characteristics that are related to handwriting neatness. For example, neat handwriting typically has clear and consistently oriented strokes, which can be reflected by HOG features. Therefore, for handwriting scoring tasks, HOG features can describe and measure the regularity and consistency of handwriting.
[0064] PCA is a dimensionality reduction method that finds a new coordinate system that maximizes the variance of all data within that new coordinate system. In other words, it identifies the main directions of variation in the data. When processing handwritten text, PCA preserves the most important data features while eliminating unnecessary details. Therefore, PCA features can reflect the overall consistency and clarity of the written text.
[0065] Wavelet transforms are used to analyze non-stationary signals or images. Applying wavelet transforms to handwritten text can extract multi-scale features such as stroke thickness, curvature, and texture. Similar to HOG, wavelet transforms can reflect the continuity and regularity of strokes. Therefore, for handwriting scoring tasks, wavelet transform features can describe and measure characteristics such as handwriting coherence and complexity, which directly affect handwriting neatness.
[0066] Step S140: performing feature extraction on each target single-word image to obtain single-word image features; the single-word image features are used to determine whether the subject has attention deficit hyperactivity disorder (ADHD).
[0067] Here, the feature extraction method includes but is not limited to histogram feature extraction method, principal component analysis, wavelet transform feature extraction method, autoencoder feature extraction method and convolutional neural network, etc.
[0068] In an embodiment of the present application, the single-word image features include at least one of the following: single-word size distribution features, vertical offset features of single words in the same row, single-word spacing features, single-word out-of-grid rate, offset features of the single-word matching grid, and the size ratio of the single word to the corresponding grid.
[0069] In the embodiments of the present application, the single-word size distribution feature is used to reflect the consistency and regularity of handwritten text size. The single-word size distribution feature is obtained by calculating the area or side length of the circumscribed rectangle of each word to obtain the distribution of handwritten text size. The distribution can be represented by statistical features such as mean, variance, maximum value, or minimum value.
[0070] In an embodiment of the present application, the vertical offset characteristics of the individual words in the same row are used to reflect the stability and coherence when writing. For the individual words in the same row, the mean of the center point of the circumscribed rectangle of these individual words or the vertical position of the lowest point of the individual words (i.e., the lowest point in the vertical direction of the individual words) can be counted, and then the deviation of the vertical position of the lowest point of each word in this row from the mean is calculated. By counting the deviation distribution of all the individual words in the full text, such as mean, variance, etc., the vertical offset characteristics of the individual words in the same row can be obtained.
[0071] In the embodiment of the present application, the single-word spacing feature is used to reflect the consistency and regularity of the single-word spacing. The single-word spacing feature is calculated by calculating the horizontal distance between the bounding rectangles of each two adjacent single words, and the text spacing distribution of the second text image can be obtained. The single-word spacing distribution can be represented by statistical features such as mean, variance, maximum or minimum value.
[0072] In the embodiments of the present application, the word out-of-grid rate is used to characterize the standardization of Chinese character writing. The bounding rectangle of each word is compared with the bounding rectangle of the grid in which the word is located to calculate the ratio of words out of grid. For example, if the bounding rectangle of a word exceeds each grid, it means that the word has been written out of each grid. The ratio of words that have been written out of grid among all the words is the word out-of-grid rate.
[0073] In the embodiment of the present application, the offset characteristics of the single-word matching grid are used to characterize the neatness of handwriting. The deviation between the center point of the circumscribed rectangle of each single word and the center point of the grid in which the word matches is calculated. The offset characteristics of all single-word matching grids can be represented using statistical features such as mean or variance.
[0074] In this embodiment of the present application, the size ratio of a character to its grid is used to reflect the relative size of a handwritten character to the standard Chinese character grid. The size ratio is calculated by calculating the average area of the bounding rectangle of each character and the average area of the grid in which the character is located; the two averages are then compared to obtain the size ratio.
[0075] In an embodiment of the present application, by extracting different features such as the size distribution characteristics of single words in the single word image, the vertical offset characteristics of the same line of single words, the single word spacing characteristics, the single word out-of-grid rate, the offset characteristics of the single word matching grid, and the size ratio of the single word to the corresponding grid, the neatness of the handwritten text of the inspected object is characterized in more detail, thereby improving the accuracy and comprehensiveness of the type of the inspected object.
[0076] In addition, to achieve better results, hardware devices such as a digital board or camera can be added to capture the writing dynamics of the inspected object and extract time-related writing motion features. Although this increases the cost, it can improve the accuracy of determining the type of the inspected object.
[0077] In the embodiments of the present application, on the one hand, only the handwritten continuous text image of the subject is needed to determine whether the subject suffers from ADHD, which eliminates the dependence on external factors such as the environment, wearable devices, and the cooperation of the subject when detecting whether the subject suffers from ADHD, saves costs, and improves applicability; on the other hand, by segmenting the continuous handwritten text image, single-word images and single-word image features are obtained, and each single-word image feature is used to represent the written text feature of the subject, thereby determining whether the subject suffers from ADHD, further improving the accuracy and comprehensiveness of determining whether the subject suffers from ADHD.
[0078] In some embodiments, reference Figure 2 As shown, the implementation of step S130 "segmenting the continuous characters in the second text image to obtain at least two target single-character images" may include the following steps S131 to S134:
[0079] Step S131: binarize the second text image to obtain a binary image.
[0080] Binarization is a process that sets the red, green, and blue (RGB) values of each pixel in the second text image to 0 (black) or 255 (white), thereby reducing the original range of color values from 256 to just two: black and white. A binary image is an image containing only two colors: black and white. The binarization method may include grayscaling the RGB color image, scanning each pixel value in the image, and setting pixel values less than a threshold to 0, and pixel values greater than or equal to the threshold to 255.
[0081] Step S132: determining at least two first connected regions in which black pixels in the binary image are connected to each other.
[0082] The first connected region refers to an image region composed of foreground pixels having the same pixel value and adjacent positions in a binary image.
[0083] Step S133: performing an expansion operation on the first connected region to obtain a second connected region.
[0084] Here, after dilating the first connected region, the blank portion in the first connected region can be filled, the image features can be enhanced, and the highlight region can be expanded, thereby obtaining the second connected region. The dilation function can be implemented by a dilate function.
[0085] Step S134: Based on the second connected area, segment the continuous text in the binary image to obtain the at least two target single-word images.
[0086] After the second connected area is obtained, the continuous characters in the binary image are segmented according to the second connected area to obtain at least two target single-character images.
[0087] Here, the segmentation processing method may include a watershed algorithm.
[0088] In the embodiment of the present application, the second text image is first binarized; a first connected region is determined; a dilation operation is performed on the first connected region; and continuous text in the binary image is processed according to the dilation operation to obtain a target single-word image. Thus, by binarizing the second text image and dilating the first connected region, the text details in the image are enhanced, thereby more accurately segmenting a complete single-word image.
[0089] In some embodiments, reference Figure 3 As shown, the implementation of step S134 "segmenting the continuous text in the binary image based on the second connected area to obtain the at least two target single-word images" may include the following steps S1341 to S1343:
[0090] Step S1341: Based on the second connected area, segment the continuous characters in the binary image to obtain at least two first candidate single-character images.
[0091] It should be understood that after segmenting continuous text, the same text may be divided into two single-word images due to errors, or punctuation marks may be segmented into single-word images. Therefore, it is necessary to screen the first candidate single-word image segmented to obtain the target single-word image.
[0092] Step S1342: Determine the size information of the circumscribed rectangle corresponding to the second connected area in each of the first candidate single-word images; the size information includes at least one of the following information: width, height, and position.
[0093] Here, the position may be the center coordinates of the circumscribed rectangle.
[0094] Step S1343: Based on the size information, the first candidate single-word images are screened to obtain the at least two target single-word images.
[0095] It should be understood that if the background information in the first text image is horizontal lines, it is only necessary to filter the first candidate single-word images to obtain the target single-word image. If the background information in the first text image is a grid, in addition to filtering the first candidate single-word images, it is also necessary to match and merge the filtered single-word images to obtain the target single-word image.
[0096] In an embodiment of the present application, the first candidate single-word image after segmentation is screened according to the size information of the circumscribed rectangle to obtain the target single-word image. In this way, the accuracy of single-word image segmentation is improved, and the accuracy of determining whether the object type suffers from ADHD is further improved.
[0097] In some embodiments, reference Figure 4 As shown, the background information in the first text image is a grid, and the implementation of step S1343 "screening the first candidate single word image based on the size information to obtain the at least two target single word images" may include the following steps S13431 to S13432:
[0098] Step S13431: remove the first candidate single-word images corresponding to the size information that does not meet the preset size information condition, and determine the remaining first candidate single-word images as second candidate single-word images.
[0099] Here, the preset size information may include a height threshold, a width threshold, or a position information deviation threshold of the circumscribed rectangle. When the size information does not meet any of the thresholds, the corresponding first candidate single word image is removed. For example, when the height or width of the size information is too small, it indicates that the first candidate single word image may contain punctuation marks or noise points, and the first candidate single word image is removed.
[0100] Step S13432: When the center coordinates of the circumscribed rectangles corresponding to at least two of the second candidate single-word images match the center coordinates of the same grid, the at least two second candidate single-word images are merged into a target single-word image to obtain the at least two target single-word images.
[0101] It should be understood that when the background information of the first text image is a grid, each single word needs to be matched with each grid to determine the grid where each single word is located, so as to better extract the single word image features later.
[0102] Here, the matching method may include: respectively calculating the distance between the center coordinates of the circumscribed rectangle corresponding to the second candidate single word image and the center coordinates of each grid, and determining the grid corresponding to the center coordinate with the smallest distance as the grid where the second candidate single word is located.
[0103] It should be noted that if the center coordinates of the circumscribed rectangles of multiple second candidate single-word images have the same minimum distance from the center coordinates of the same grid, the multiple second candidate single-word images are merged into one target single-word image.
[0104] In an embodiment of the present application, after the first candidate single-word image after segmentation is screened according to the size information of the circumscribed rectangle, the center coordinates of the circumscribed rectangle are further matched and merged with the center coordinates of the grid to determine the grid where the candidate single word is located, and the target single-word image is obtained. In this way, the accuracy of single-word image segmentation is improved, and the accuracy of object type is further improved; and, by performing single-word segmentation and obtaining the circumscribed rectangle, the position and range of each single word can be obtained, so as to facilitate better extraction of single-word image features in the later stage, provide convenience for subsequent neatness analysis, and further improve the accuracy of whether the object type suffers from ADHD.
[0105] In addition, this application only extracts single image features related to handwriting neatness. If optical character recognition (OCR) is performed on handwritten text, features related to semantic errors, expression logic, etc. can be extracted to improve the accuracy of determining the type of the inspected object.
[0106] In some embodiments, the image processing method may further include the following steps S210 to S240:
[0107] Step S210: Determine edge information in the first text image.
[0108] Here, the first text image is first converted to grayscale; and then edge detection is used to extract edge information of the first text image.
[0109] Step S220: using the Hough transform algorithm to determine the straight line expression corresponding to the edge information.
[0110] Here, the line group expression is a line group expression that is continuous, has a length exceeding a certain threshold, and has uniform spacing, and can be a horizontal line group expression and / or a vertical line group expression.
[0111] The Hough transform algorithm is used to detect shapes such as straight lines and circles in an image. The basic idea is to map points on the first text image plane to a parameter space, and by voting and accumulating in the parameter space, find the peak of the parameter space, that is, detect the desired shape. Generally, the Hough transform is performed using the OpenCV v2.HoughLines function to obtain the line group expression. Only expressions for continuous line groups with lengths exceeding a certain threshold, horizontal or vertical angles, and uniform spacing are stored;
[0112] Step S230: When the straight line expression only includes a horizontal straight line expression, determining that the background information in the first text image is the horizontal line.
[0113] Step S240: When the line group expression includes a horizontal line expression and a vertical line expression, determining that the background information in the first text image is the grid.
[0114] It should be understood that if only a horizontal straight line expression exists in the edge information, it indicates that the background information is a horizontal line; if both a horizontal straight line expression and a vertical straight line expression exist in the edge information, it indicates that the background information is a grid.
[0115] In another embodiment of the present application, if the background information in the first text image is detected to be a grid, before deleting the grid background, the corner coordinates of all grids in the first text image are calculated based on the expression of each line; then, based on the corner coordinates, the length, width, and center coordinates of each grid are calculated. The corner coordinates refer to the vertex coordinates of the grid. Here, the center coordinates of the grid can be used to match the center coordinates of the circumscribed rectangle of the second candidate single-word image.
[0116] In an embodiment of the present application, the background information of the first text image is identified and then removed, and only the Chinese character image is retained, thereby preventing the background from affecting the accuracy of the screening result of whether the object type suffers from ADHD.
[0117] In some embodiments, the implementation of step S120 of "removing background information from the first text image to obtain a second text image" may include the following steps S121 to S123:
[0118] Step S121: creating a first mask for each straight line in the background information; wherein the first mask is a white line that obeys the straight line expression of each straight line.
[0119] The first mask is mainly used to block each line in the grid or horizontal line in the background information. It should be understood that when the background information is a horizontal line, only horizontal line masks need to be created; when the background information is a grid, horizontal line masks and vertical line masks need to be created.
[0120] Step S122: performing an expansion operation on the white lines in the first mask to obtain a second mask.
[0121] Here, the dilation operation is used to increase the width of the white line until it is no less than the thickness of the grid or horizontal line in the background information.
[0122] Step S123: merging the second mask with the first text image to obtain the second text image.
[0123] The merging process refers to applying the second mask to each straight line in the background information so that the second mask covers each straight line in the background information, thereby obtaining the first text image.
[0124] It should be noted that when using a mask to remove the background, a small amount of text may be erased. A repair algorithm (such as the cv2.inpaint function of OpenCV) can be used to repair the erased text part to obtain the first text image.
[0125] In the embodiment of the present application, after determining that there are grids or horizontal lines in the image, the grids or horizontal lines are removed using a mask to obtain an image of pure handwritten text content, which helps to improve the accuracy of subsequent single-word separation.
[0126] In the embodiment of the present application, the method further includes the following steps 310 to 320:
[0127] Step 310: Obtain text information corresponding to the subject; the text information includes at least one of the following information: age, gender, name, and dominant hand.
[0128] Here, the subjects can be people of all ages. Subject information can include age, gender, name, and handedness. Handedness can also be used as a feature of the subject to assess their handwriting. For example, left-handed and right-handed handwriting tend to tilt differently.
[0129] Step 320: Input the single-word image features and corresponding text information corresponding to the subject into a pre-trained classification model, and output a screening result for ADHD of the subject; the pre-trained classification model is used to determine the screening result of the subject.
[0130] In the embodiment of the present application, the pre-trained classification model is a pre-trained image processing model. Here, the pre-trained classification model can be obtained by the electronic device through training on sample data, or it can be obtained by the electronic device from other servers providing models.
[0131] In an embodiment of the present application, the text information corresponding to the subject and the corresponding single-word image features are input into a pre-trained classification model to quickly output the result of whether the subject suffers from ADHD, thereby improving the accuracy and comprehensiveness of screening whether the subject suffers from ADHD.
[0132] In some embodiments, the training process of the classification model includes the following steps S321 to S324:
[0133] Step S321: Acquire the single-word image features and the sample label of the sample image; the sample label includes the result of whether the subject corresponding to the sample image has ADHD.
[0134] Here, the single-word image features and sample labels of the sample images can be understood as a training data set. The sample labels can be manually annotated or obtained by other means. The method for obtaining the single-word image features can refer to the above steps S110-140 and will not be repeated here.
[0135] In addition, the text features of the subject information (such as age, gender and handedness information) can be converted into digital codes or binary codes and added to the feature vector of the single-word image features to increase the predictive ability and accuracy of the training model.
[0136] In the embodiment of the present application, the classification model can be a binary classification model or a multi-classification model. For example, in a binary classification model, the sample labels include yes and no.
[0137] Step S322: Processing the single-word image feature based on the classification model to be trained to obtain a first output result; the first output result is used to indicate whether the subject corresponding to the single-word image feature suffers from ADHD.
[0138] Here, the acquired single-word image features may be normalized or standardized to avoid excessive scale differences between different features and ensure comparability between the features.
[0139] In the embodiment of the present application, the classification model to be trained can be a selected classifier, such as an extreme gradient boosting algorithm (XGBoost), a logistic regression (LR), a support vector machine (SVM), or a neural network classification algorithm model. The embodiment of the present application is not limited here.
[0140] Step S323: Determine a first difference value between the sample label and the first output result through a target loss function.
[0141] In this embodiment of the present application, since the classification model to be trained is the initial model, the predicted subject type in the first output result obtained after the sample data is processed by the classification model to be trained is not the actual type of the subject being tested. The sample label is the actual result of whether the subject has ADHD. Therefore, there is a difference between the sample label and the first output result.
[0142] Here, the difference between the sample label and the first output result can be quantified using a target loss function. In other words, the target loss function can be a function that measures the degree of inconsistency between the predicted value of the recognition model to be trained and the true value. Generally, the greater the difference between the sample label and the first output result, the larger the calculated first difference value. This first difference value can, to a certain extent, characterize the classification performance of the classification model to be trained.
[0143] Step S324: training the classification model to be trained based on the first difference value until a training end condition is met, thereby obtaining the pre-trained classification model.
[0144] In the embodiment provided in the present application, after the first difference value is determined, the classification model to be trained can be iteratively trained based on the first difference value and sample features, and finally a pre-trained classification model is obtained through training.
[0145] The training process of the pre-trained classification model is achieved through iterative training. The iterative process continuously adjusts the model parameters using the first difference value, the second difference value, and so on until the adjusted classification model to be trained converges or the obtained Nth difference value is within the preset threshold range, determining that the training end condition has been met. Here, N is an integer greater than or equal to 1. The adjusted classification model to be trained at this point is the final pre-trained classification model.
[0146] In an embodiment of the present application, during the training process of the classification model to be trained, the parameters in the classification model to be trained that determine whether the subject suffers from ADHD are trained together, which can reduce the computational complexity of image processing and improve the training efficiency and speed; and the trained classification model can be used directly to evaluate whether the subject suffers from ADHD, thereby improving the speed and efficiency of detection.
[0147] The image processing method provided in the embodiment of the present application is described in detail below in conjunction with specific application scenarios.
[0148] It should be understood that handwriting, as a fine motor activity involving visual-motor integration, spatial perception, cognition, and attention, reflects a wide range of information. Many studies have shown that, compared with healthy controls, individuals with ADHD have lower control over their handwriting movements, less fluent handwriting, longer delays between instruction and handwriting initiation, smaller planned movement amplitudes, and less fluent handwriting. These differences have all reached significant levels.
[0149] For patients with ADHD, the onset of the disease often begins in school or preschool, and composition is an important medical history document that directly reflects the patient's writing level. Chinese composition writing is closely related to visual sensory integration, motor planning, and attention levels. Compared with the normal group, due to defects in the brain's attention impulse control, ADHD patients tend to have untidy and sloppy handwriting. Collecting images of written Chinese composition and conducting an objective and quantitative analysis of their handwriting neatness as an early screening method is easy to popularize, highly objective, and highly sensitive.
[0150] The application process of the image processing method provided in the embodiment of the present application in screening ADHD is described in detail below. Figure 5 As shown, the process includes the following steps S51 to S511:
[0151] Step S51: input a third text image.
[0152] The third text image is a child's handwritten composition image, wherein the handwritten composition image includes continuous Chinese characters.
[0153] Step S52: performing enhancement processing on the third text image.
[0154] In the embodiment of the present application, enhancement processing can be understood as preprocessing, which may include histogram equalization, filtering, and image enhancement.
[0155] Among them, histogram equalization is used to eliminate the uneven illumination phenomenon caused by inconsistent illumination intensity in different areas of the third text image. The algorithm enhances the overall brightness and contrast of the image by mapping the pixel values of the third text image to a uniform distribution, making the illumination of the area where the handwritten text is located more uniform and improving the uniformity of the text area.
[0156] Filtering is used to reduce the impact of random interference points introduced by various factors on an image. Common filtering methods include Gaussian filtering and median filtering. Gaussian filtering smoothes an image by taking a weighted average of the neighborhood surrounding each pixel. Median filtering replaces the current pixel value with the median value of all pixels in the neighborhood, effectively removing noise points.
[0157] Image enhancement aims to improve the quality of images and the saliency of text. Generally, image enhancement can apply a sharpening filter to enhance edge information, or use an adaptive threshold segmentation algorithm to binarize the image to highlight the text information area and suppress background noise.
[0158] Step S53: performing correction processing on the enhanced third text image to obtain the first text image.
[0159] It should be understood that the correction process may include tilt correction. Generally, the correction process method may include but is not limited to: Hough line detection, projection (Radon) angle transformation and other methods.
[0160] Exemplarily, the implementation of the correction process may include: first, converting the enhanced third text image into a grayscale image; second, using edge detection (such as Canny operator edge detection) to find the edges of the image and obtain a binary edge image; third, using a contour extraction algorithm (such as OpenCV's cv2.findContours function) to extract all contours in the binary edge image; fourth, finding the quadrilateral contour with the largest area among all contours; fifth, extracting the coordinates of the four corner points of the quadrilateral contour; sixth, calling the perspective transformation function (such as OpenCV's cv2.getPerspectiveTransform function) to calculate the perspective transformation matrix for correcting the tilt through the corner point coordinates; seventh, applying the perspective transformation matrix to correct the enhanced handwritten composition image; eighth, calculating the corrected corner point coordinates, and cutting out the rectangular area formed by the corner points as the handwritten text area to obtain the first text image.
[0161] Step S54: Identify whether there is a background grid or horizontal line in the first text image. If so, execute step S55; otherwise, execute step S56.
[0162] Handwritten essays often have small grids or lines on the paper to standardize the writing area or facilitate word counting. When analyzing the neatness of handwritten text, these background lines need to be removed to avoid affecting the accuracy of ADHD screening.
[0163] Here, the implementation of identifying grids or horizontal lines can include: first, graying the corrected composition image and performing edge detection to obtain edge information; then using OpenCV's cv2.HoughLines function to perform Hough line transform on the edge information, and only storing continuous straight line expressions with a length exceeding a certain threshold, horizontal or vertical angles, and uniform spacing; then detecting the direction of the straight line expression, if the detected straight line group is empty, it proves that the composition image background is blank; if only a horizontal straight line group is detected, it proves that the composition image background is a horizontal line; if both horizontal and vertical straight line groups are detected, it proves that the composition image background is a grid; if the composition image background is a grid, then according to the expression of each straight line in the straight line group, the coordinates of the corner points of all text grids in the image area are calculated; then according to the corner point coordinates, the length, width and center point coordinates of each grid are calculated and recorded.
[0164] Step S55: removing the background grid or horizontal lines in the first image.
[0165] Step S56: Obtain a second text image.
[0166] Here, the second text image is a composition image without background, wherein a mask is generally used to remove grids or horizontal lines.
[0167] Step S57: Segment the second text image to obtain single-word images.
[0168] Here, the segmentation processing methods include but are not limited to: segmentation based on connected regions, segmentation based on projection, and segmentation based on machine learning (or deep learning).
[0169] Step S58: extracting features from the single-word image to obtain single-word image features.
[0170] It should be understood that the handwritten text of ADHD patients is more sloppy and untidy than that of healthy people. Therefore, after completing the segmentation of single words in the handwritten text in the image, it is necessary to further extract various quantitative features of handwriting neatness to characterize the difference between ADHD patients and healthy people.
[0171] Among them, the single-word image features may include: single-word size distribution features, vertical offset features of single words in the same row, single-word spacing features, single-word out-of-grid rate, offset features of single-word matching grids, and the size ratio of single words to corresponding grids.
[0172] Step S59: extracting global writing features from the second text image.
[0173] It should be understood that feature extraction can be performed directly on the second text image (ie, the handwritten composition image without background), and the writing features at the full text level can be used to further quantitatively analyze and evaluate the neatness of the handwritten text.
[0174] Among them, image feature methods for extracting full-text writing features may include: HOG features, PCA features, wavelet transform features and other methods.
[0175] Step S510: inputting the global writing features, single-word image features and text information of the subject into a pre-trained classification model.
[0176] It should be understood that the text information of the subject includes the child's age, gender, and handedness (e.g., left-handed, right-handed). Age and gender are important indicators for determining the baseline writing of the group to which the subject belongs; handedness can help evaluate the text characteristics of the subject's handwriting.
[0177] Here, the global writing features, the single word image features and the information of the inspected object are in one-to-one correspondence. The global writing features, the single word image features and the information of the inspected object are converted into feature vectors and input into the pre-trained classification model.
[0178] Here, the preselected classification model may be an ADHD screening model. In essence, the ADHD screening model is a binary classification model for outputting whether the subject suffers from ADHD.
[0179] Step S511: Output the screening result of the subject.
[0180] Here, the screening result of the subject is whether or not the subject has ADHD.
[0181] The image processing method provided in the embodiment of the present application, on the one hand, includes processing such as tilt correction and background removal of the image, extracting the writing features of single words and full texts, while taking into account non-writing features such as gender, age, and handedness, integrating all collected information, and finally establishing an objective and quantitative ADHD screening model based on handwriting neatness; on the other hand, it includes statistical characteristics of the ratio of words written out of the grid, statistical characteristics of the distribution of word size, statistical characteristics of the offset of words from the center of the grid, statistical characteristics of the upper and lower offsets of words in the same line, statistical characteristics of the distribution of word spacing, etc.
[0182] The use of the image processing method provided by the embodiments of this application in ADHD application scenarios has the following advantages over existing technologies: Passive monitoring is easy to operate and highly acceptable: no active participation or cooperation is required from the subject; only images of the subject's writing in their daily schoolwork are needed. No specific tasks or stimuli are required, and handwriting samples can be obtained from the subject's natural environment. Compared with active screening methods, the subject's compliance and acceptance are higher; no wearable devices are required, and the method is highly universal and applicable; High sensitivity to daily behavior: The collected handwritten writing images of the subject are data generated in daily life, which can directly reflect the subject's behavioral state in daily life. Compared with screening conducted in a controlled test environment, it provides an assessment that is more in line with real-life situations, enhancing the sensitivity and specificity of the screening; Digital and intelligent automatic screening, with high popularity: The fully automated screening solution based on image processing and machine learning algorithms does not require the full guidance of professional physicians, which can increase the popularity of ADHD screening; Improved screening accuracy and comprehensiveness: The evaluation of the neatness of writing can reflect the daily behavioral function of the subject's brain attention impulse control mechanism. Existing methods primarily focus on acquiring physiological signals, while handwriting neatness provides a more behavioral and visual assessment method. By leveraging additional behavioral indicators, the image processing method provided in this application may provide a more comprehensive and multi-faceted ADHD screening approach, thereby complementing existing physiological indicators for ADHD screening and improving the accuracy and comprehensiveness of screening.
[0183] The present application also provides an image processing device, such as Figure 6 As shown, the image processing device 600 includes:
[0184] An acquisition module 601 is configured to acquire a first text image containing continuous characters handwritten by a subject;
[0185] A removal module 602 is configured to remove background information from the first text image to obtain a second text image; the background information includes: a grid or horizontal lines;
[0186] a segmentation module 603, configured to segment the continuous characters in the second text image to obtain at least two target single-character images;
[0187] The extraction module 604 is used to extract features from each target single-word image to obtain single-word image features; the single-word image features are used to determine whether the subject has attention deficit hyperactivity disorder (ADHD).
[0188] In some embodiments, the segmentation module 603 includes: a binarization processing submodule, used to binarize the second text image to obtain a binary image; a first determination submodule, used to determine at least two first connected areas in which black pixels in the binary image are connected to each other; a first expansion submodule, used to perform an expansion operation on the first connected area to obtain a second connected area; and a first segmentation submodule, used to segment the continuous text in the binary image based on the second connected area to obtain the at least two target single-word images.
[0189] In some embodiments, the first segmentation submodule includes: a first segmentation unit, used to segment the continuous text in the binary image based on the second connected area to obtain at least two first candidate single-word images; a first determination unit, used to determine the size information of the circumscribed rectangle corresponding to the second connected area in each of the first candidate single-word images; the size information includes at least one of the following information: width, height and position; a screening unit, used to screen the first candidate single-word images based on the size information to obtain the at least two target single-word images.
[0190] In some embodiments, the background information in the first text image is a grid; the screening unit includes: a removal subunit, used to remove the first candidate single-word images corresponding to the size information that does not meet the preset size information condition, and determine the remaining first candidate single-word images as second candidate single-word images; a merging subunit, used to merge the at least two second candidate single-word images into a target single-word image when the center coordinates of the circumscribed rectangles corresponding to at least two of the second candidate single-word images match the center coordinates of the same grid, thereby obtaining the at least two target single-word images.
[0191] In some embodiments, the image processing device 600 also includes: a first determination module for determining edge information in the first text image; a second determination module for determining a straight line expression corresponding to the edge information using a Hough transform algorithm; a third determination module for determining that the background information in the first text image is the horizontal line when the straight line expression only includes a horizontal straight line group expression; and a fourth determination module for determining that the background information in the first text image is the grid when the straight line group expression includes a horizontal straight line expression and a vertical straight line group expression.
[0192] In some embodiments, the removal module 602 includes: a creation submodule for creating a first mask for each straight line in the background information; wherein the first mask is a white line that obeys the straight line expression of each straight line; a second expansion submodule for performing an expansion operation on the white lines in the first mask to obtain a second mask; and a merging submodule for merging the second mask with the first text image to obtain the second text image.
[0193] In some embodiments, the image processing device 600 also includes: a third acquisition module, used to obtain text information corresponding to the subject; the text information includes at least one of the following information: age, gender, name, and dominant hand; an input module, used to input the single-word image features and corresponding text information corresponding to the subject into a pre-trained classification model, and output a screening result for ADHD of the subject; the pre-trained classification model is used to determine the screening result of the subject.
[0194] In some embodiments, the input module also includes: an acquisition submodule, used to obtain the single-word image features and the sample label of the sample image; the sample label includes the result of whether the subject corresponding to the sample image has ADHD; an output submodule, used to process the single-word image features based on the classification model to be trained to obtain a first output result; the first output result is used to characterize whether the subject corresponding to the single-word image features suffers from ADHD; a second determination submodule, used to determine a first difference value between the sample label and the first output result through a target loss function; a training submodule, used to train the classification model to be trained based on the first difference value until the training end condition is met, thereby obtaining the pre-trained classification model.
[0195] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention is shown in FIG. Figure 7 As shown, in actual application, the electronic device 700 includes: a memory 701 for storing computer executable instructions; a processor 702, connected to the memory 701, for implementing the position determination method by executing the computer executable instructions; a bus system 703 for coupling the various components together. It can be understood that the bus system 703 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 703 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 7 The various buses are labeled bus system 703 .
[0196] It is understood that the memory in this embodiment can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.
[0197] The methods disclosed in the above embodiments of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above methods can be completed by hardware integrated logic circuits in the processor or instructions in software form. The above processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory. The processor reads the information in the memory and completes the steps of the above methods in combination with its hardware.
[0198] The present application also provides a computer storage medium, specifically a computer-readable storage medium, on which computer instructions are stored. As a first embodiment, when the computer storage medium is located in an electronic device, the computer instructions, when executed by a processor, implement any step of the above-mentioned image processing method of the present application.
[0199] An embodiment of the present application provides a computer program product containing instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0200] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0201] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0202] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or at least two units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0203] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, ROM, RAM, disks or optical disks, etc. Various media that can store program codes.
[0204] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0205] It should be noted that the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0206] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a first text image containing continuous characters written by the subject; removing background information from the first text image to obtain a second text image; the background information includes: a grid or horizontal lines; Segmenting the continuous characters in the second text image to obtain at least two target single-character images; Feature extraction is performed on each target single-word image to obtain single-word image features; the single-word image features are used to determine whether the subject has attention deficit hyperactivity disorder (ADHD).
2. The method according to claim 1, characterized in that The segmenting process of the continuous characters in the second text image to obtain at least two target single-character images includes: performing binarization processing on the second text image to obtain a binary image; Determine at least two first connected regions in the binary image where black pixels are connected to each other; performing an expansion operation on the first connected region to obtain a second connected region; Based on the second connected area, the continuous characters in the binary image are segmented to obtain the at least two target single-character images.
3. The method according to claim 2, characterized in that Based on the second connected region, segmenting the continuous characters in the binary image to obtain the at least two target single-character images includes: Segmenting the continuous characters in the binary image based on the second connected region to obtain at least two first candidate single-character images; Determine size information of a circumscribed rectangle corresponding to the second connected region in each of the first candidate single-word images; the size information includes at least one of the following information: width, height, and position; Based on the size information, the first candidate single-word images are screened to obtain the at least two target single-word images.
4. The method according to claim 3, characterized in that The background information in the first text image is a grid; The step of screening the first candidate single-word image based on the size information to obtain the target single-word image includes: Remove the first candidate single word images whose size information does not meet the preset size information condition, and determine the remaining first candidate single word images as second candidate single word images; When the center coordinates of the circumscribed rectangles corresponding to at least two of the second candidate single-word images match the center coordinates of the same grid, the at least two second candidate single-word images are merged into a target single-word image to obtain the at least two target single-word images.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: determining edge information in the first text image; Using the Hough transform algorithm, determining a straight line expression corresponding to the edge information; In a case where the straight line expression only includes a horizontal straight line group expression, determining that the background information in the first text image is the horizontal line; In a case where the straight line group expression includes a horizontal straight line expression and a vertical straight line group expression, the background information in the first text image is determined to be the grid.
6. The method according to any one of claims 1 to 4, characterized in that Removing background information from the first text image to obtain a second text image includes: Creating a first mask for each straight line in the background information; wherein the first mask is a white line that obeys the straight line expression of each straight line; Performing a dilation operation on the white lines in the first mask to obtain a second mask; The second mask is merged with the first text image to obtain the second text image.
7. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Obtaining text information corresponding to the subject; the text information includes at least one of the following information: age, gender, name, and dominant hand; The single-word image features and corresponding text information corresponding to the subject are input into a pre-trained classification model, and a screening result for ADHD of the subject is output; the pre-trained classification model is used to determine the screening result of the subject.
8. The method according to claim 7, characterized in that The training process of the pre-trained classification model includes: Obtaining a single-word image feature and the sample label of a sample image; the sample label includes a result of whether the subject corresponding to the sample image has ADHD; Processing the single-word image feature based on the classification model to be trained to obtain a first output result; the first output result is used to indicate whether the subject corresponding to the single-word image feature suffers from ADHD; Determining a first difference value between the sample label and the first output result through a target loss function; The classification model to be trained is trained based on the first difference value until a training end condition is met, thereby obtaining the pre-trained classification model.
9. An image processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a first text image containing continuous characters handwritten by the subject; a removal module, configured to remove background information from the first text image to obtain a second text image; the background information includes: a grid or horizontal lines; a segmentation module, configured to segment the continuous characters in the second text image to obtain at least two target single-character images; The extraction module is used to extract features from each target single-word image to obtain single-word image features; the single-word image features are used to determine whether the subject has attention deficit hyperactivity disorder (ADHD).
10. An electronic device, characterized in that: include: a memory for storing computer-executable instructions; A processor is connected to the memory, and is used to implement the image processing method according to any one of claims 1 to 8 by executing the computer-executable instructions.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, the image processing method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that The method comprises computer program instructions, which implement the image processing method according to any one of claims 1 to 8 when executed by a processor.