Method and device for identifying mild cognitive impairment of old people based on image identification
By acquiring image sequences of elderly people through image recognition technology and generating object description information, the problem of difficulty in timely identification of mild cognitive impairment has been solved, enabling effective identification and early warning of mild cognitive impairment and reducing the risk of Alzheimer's disease.
Patent Information
- Application Number
- CN202511168383.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-20
AI Technical Summary
In existing technologies, the onset of mild cognitive impairment is a long-term process, and conventional medical diagnostic methods are unable to identify changes in mild cognitive impairment in a timely manner, increasing the risk of developing Alzheimer's disease and other conditions.
By using image recognition methods, image sequences of elderly people are acquired, object recognition and image extraction are performed, and object description information is generated, including predicted action types, identified action types and emotional description information. Combining behavioral and emotional changes, mild cognitive impairment assessment information is generated and synchronized to the smart elderly care platform.
It enables the effective identification of mild cognitive impairment, timely detection of changes in patients, and reduces the risk of developing Alzheimer's disease and other conditions.
Smart Images

Figure CN120998468A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of computer technology, the field of mild cognitive impairment evaluation, and in particular to a method and device for identifying mild cognitive impairment of the elderly based on image recognition. BACKGROUND
[0002] Mild cognitive impairment (MCI) refers to the decline of memory or other cognitive functions, but does not affect the basic daily life, which is between the normal state and dementia. When a patient has mild cognitive impairment, it shows the decline of one or more cognitive functions, but no obvious dementia. According to incomplete statistics, the prevalence of mild cognitive impairment in the elderly is high, and it has the risk of developing into Alzheimer's disease and the like. At present, mild cognitive impairment is mainly determined by medical diagnosis.
[0003] However, the onset process of mild cognitive impairment has the characteristics of long periodicity, and the regular medical diagnosis method for regular follow-up diagnosis has poor recognition effect on mild cognitive impairment in the window period, which makes it difficult to find the changes of mild cognitive impairment of patients in time, thereby increasing the risk of developing into Alzheimer's disease and the like. SUMMARY
[0004] The summary of the present disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments part. The summary of the present disclosure does not aim to identify the key features or essential features of the claimed technical solutions, nor does it aim to limit the scope of the claimed technical solutions.
[0005] Some embodiments of the present disclosure propose a method and device for identifying mild cognitive impairment of the elderly based on image recognition to solve the technical problems mentioned in the background part.
[0006] In a first aspect, some embodiments of the present disclosure provide an image recognition-based mild cognitive impairment identification method for the elderly. The method is applied to a smart elderly care platform and includes: obtaining an initial image sequence for a to-be-identified object, wherein the to-be-identified object is an object corresponding to an age greater than a preset age and pre-labeled to be subjected to mild cognitive impairment identification; performing object identification on an initial image in the initial image sequence to obtain a target position sequence, wherein the target position represents the sequence position of the initial image containing the to-be-identified object in the initial image sequence; performing image extraction on the initial image sequence according to the target position sequence to obtain a target image sequence, wherein the target image corresponds to the target position; generating an object description information set corresponding to the to-be-identified object according to a pre-trained object description information generation model and the target image sequence, wherein the object description information includes a predicted action type, an identified action type, and emotion description information, and the emotion description information represents the emotional changes of the to-be-identified object; generating mild cognitive impairment evaluation information for the to-be-identified object according to the object description information set; and synchronizing the mild cognitive impairment evaluation information to the smart elderly care platform.
[0007] In a second aspect, some embodiments of the present disclosure provide an image recognition-based mild cognitive impairment identification device for the elderly. The device includes: an obtaining unit configured to obtain an initial image sequence for a to-be-identified object, wherein the to-be-identified object is an object corresponding to an age greater than a preset age and pre-labeled to be subjected to mild cognitive impairment identification; an object identification unit configured to perform object identification on an initial image in the initial image sequence to obtain a target position sequence, wherein the target position represents the sequence position of the initial image containing the to-be-identified object in the initial image sequence; an image extraction unit configured to perform image extraction on the initial image sequence according to the target position sequence to obtain a target image sequence, wherein the target image corresponds to the target position; a first generation unit configured to generate an object description information set corresponding to the to-be-identified object according to a pre-trained object description information generation model and the target image sequence, wherein the object description information includes a predicted action type, an identified action type, and emotion description information, and the emotion description information represents the emotional changes of the to-be-identified object; a second generation unit configured to generate mild cognitive impairment evaluation information for the to-be-identified object according to the object description information set; and a synchronization unit configured to synchronize the mild cognitive impairment evaluation information to the smart elderly care platform.
[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device. The electronic device includes one or more processors and a storage device having one or more programs stored thereon. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementation manners of the first aspect.
[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0010] The above-described embodiments of this disclosure have the following beneficial effects: The image recognition-based method for identifying mild cognitive impairment in the elderly, as described in some embodiments of this disclosure, effectively identifies changes in mild cognitive impairment in patients. Specifically, the reason for the inability to identify it in a timely manner is that the onset of mild cognitive impairment is characterized by a long period. Conventional medical diagnostic methods involving regular follow-up diagnoses are ineffective in identifying mild cognitive impairment within the window period, making it difficult to detect changes in the patient's mild cognitive impairment in a timely manner, thereby increasing the risk of developing Alzheimer's disease, etc. For example, assuming a patient needs to undergo medical diagnosis at time points T1 and T2, if mild cognitive impairment changes occur within the time window formed by T1 and T2, it is difficult to detect in a timely manner, especially as the window period increases, further increasing the risk of deterioration. Based on this, the image recognition-based method for identifying mild cognitive impairment in the elderly, as described in some embodiments of this disclosure, firstly acquires an initial image sequence for the object to be identified, wherein the object to be identified is a pre-marked object whose age is greater than a preset age and is intended for mild cognitive impairment identification. Secondly, object recognition is performed on the initial images in the aforementioned initial image sequence to obtain a target location sequence, where the target location represents the sequence position of the initial image containing the object to be identified within the initial image sequence. Next, based on the target location sequence, image extraction is performed on the initial image sequence to obtain a target image sequence, where each target image corresponds to a target location. In practice, the object to be identified, as the subject of identification for mild cognitive impairment, is not static; therefore, there may be situations where the object to be identified is not present in the image. Since the onset of mild cognitive impairment is characterized by a long period, a large number of images need to be processed for identification. To avoid unnecessary computational resource consumption, object recognition is used to determine the target location, thus filtering out invalid image frames. Furthermore, based on a pre-trained object description information generation model and the aforementioned target image sequence, a set of object description information corresponding to the object to be identified is generated. This object description information includes: predicted action type, identified action type, and emotion description information, where the emotion description information represents the emotional changes of the object to be identified. In practice, patients with mild cognitive impairment often experience a decline in one or more cognitive functions. Therefore, this disclosure generates a set of object description information (which changes over time) for the identified individual, based on both behavioral and emotional dimensions. Next, based on this object description information set, mild cognitive impairment assessment information for the identified individual is generated. This automatically generates a mild cognitive impairment assessment for the identified individual from both behavioral and emotional dimensions. Finally, the mild cognitive impairment assessment information is synchronized to the aforementioned smart elderly care platform. This method enables the effective identification of changes in mild cognitive impairment in patients. Attached Figure Description
[0011] The above and other features, aspects and advantages of the present disclosure will become more apparent after a reading of the following detailed description together with the accompanying drawings. Throughout the drawings, similar or same reference numerals are used to denote similar or same elements. It should be understood that the drawings are schematic and elements and features are not necessarily to scale.
[0012] Figure 1 is a flowchart of some embodiments of the image recognition based mild cognitive impairment identification method for the elderly according to the present disclosure;
[0013] Figure 2 is a schematic diagram of the generation process of object description information;
[0014] Figure 3 is a schematic diagram of the structure of some embodiments of the image recognition based mild cognitive impairment identification device for the elderly according to the present disclosure;
[0015] Figure 4 is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0017] It should also be noted that, for the sake of brevity, only the parts of the drawings that are relevant to the present disclosure are shown. The embodiments and features in the present disclosure can be combined with each other in the case of no conflict.
[0018] It should be noted that the terms “first”, “second”, and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0019] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly stated in the context.
[0020] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of the messages or information.
[0021] The collection, storage, use, etc. of the initial image sequence of the to-be-identified object involved in the present disclosure have been performed by the relevant organization or individual before the corresponding operation is performed, and the relevant organization or individual has fulfilled obligations including but not limited to: conducting a personal information security impact assessment, fulfilling the obligation of informing the principal of personal information, and obtaining the prior authorization consent of the principal of personal information. In particular, since the to-be-identified object may have one or more cognitive function declines, the relevant organization or individual has also fulfilled the above obligations to its legal guardian.
[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0023] Reference Figure 1 Fig. 1 shows a flow 100 of some embodiments of the image recognition-based mild cognitive impairment identification method for the elderly according to the present disclosure. The image recognition-based mild cognitive impairment identification method for the elderly includes the following steps:
[0024] Step 101, obtaining an initial image sequence for a to-be-identified object.
[0025] In some embodiments, the subject (for example, a computing device) of the execution of the image recognition-based mild cognitive impairment identification method for the elderly can obtain the initial image sequence for the to-be-identified object through wired or wireless connection.
[0026] Wherein, the to-be-identified object is an object corresponding to an age greater than a preset age, which is marked in advance for mild cognitive impairment identification. Specifically, the preset age can be set according to the statistical incidence of mild cognitive impairment in different age groups. For example, there is a survey showing that the incidence of mild cognitive impairment in the elderly is 25% to 56%, so the preset age can be set to 50 years old. Correspondingly, in order to improve the coverage of mild cognitive impairment identification, the preset age can also be lowered.
[0027] Wherein, the collection of the initial image sequence can be collected by the camera arranged in the activity area of the to-be-identified object.
[0028] Example 1, taking a to-be-identified object as an example, a single elderly person, a plurality of cameras can be arranged in the residence of the to-be-identified object for collecting the initial image sequence.
[0029] In Example 2, taking the non-living-alone old as an example, a plurality of cameras can be arranged in the activity area (e.g., residence) of the non-living-alone old to collect the initial image sequence. In particular, since the activity area contains not only the non-living-alone old, in order to ensure the effectiveness and pertinence of the collection, the non-living-alone old can wear a positioning device (e.g., positioning bracelet, etc.), wherein the positioning device is embedded with an Ultra Wide Band (UWB) transmitter. The camera can be provided with at least two UWB receivers. The UWB receiver and the UWB transmitter can interact with each other. The time difference between the signals received by the two UWB receivers is used to locate the non-living-alone old, so as to control the camera to perform operations including but not limited to rotation, zooming, etc., so as to ensure that the initial image contains only the non-living-alone old as much as possible, and when the non-living-alone old appears in the collection range of the camera, the non-living-alone old appears in the picture (in the initial image) as much as possible.
[0030] In practice, for example, the computing device can be deployed in the activity area corresponding to the non-living-alone old, at this time, the computing device can acquire the initial image sequence from the camera through wired or wireless connection, and the computing device can be deployed in the activity area to achieve the purpose of local data processing. For another example, the computing device can also be deployed remotely, at this time, the computing device can acquire the initial image sequence from the camera through wireless connection. In particular, during the transmission of the initial image sequence, the initial image needs to be encrypted to ensure the security of data transmission.
[0031] It should be noted that the wireless connection mode can include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (Ultra Wide Band) connection, and other now known or future developed wireless connection modes.
[0032] It should be noted that the computing device can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as a plurality of software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made herein.
[0033] In step 102, the initial images in the initial image sequence are subjected to object recognition to obtain a target position sequence.
[0034] In some embodiments, the execution subject can perform object recognition on the initial images in the initial image sequence to obtain a target position sequence.
[0035] The target position represents a sequence position of the initial image containing the object to be recognized in the initial image sequence.
[0036] In practice, a target detection model such as Tiny-YOLO (You Only Look Once) can be used to perform object recognition on the initial images in the initial image sequence to obtain the target position sequence.
[0037] In some optional implementations of some embodiments, the execution subject performs object recognition on the initial images in the initial image sequence to obtain the target position sequence, including:
[0038] Step S1: determining whether there is a camera angle change of the target camera in a target time period.
[0039] The target camera is the camera that collects the initial image sequence, and the target time period is the collection time period corresponding to the initial image sequence. The target camera can be provided with an IMU (Inertial Measurement Unit) module, so the camera angle change of the target camera can be monitored by the IMU module.
[0040] In practice, first, when the target camera has an angle change, a record composed of the angle change value monitored by the IMU module and the corresponding timestamp can be generated and stored. Then, whether there is a camera angle change of the target camera in the target time period can be determined by judging whether the angle change values contained in the records with timestamps in the target time period are consistent.
[0041] Step S2: in response to the absence of a camera angle change p, obtaining a target background image.
[0042] The target background image is a static background image under the current camera angle corresponding to the target camera.
[0043] In practice, the target camera will collect and cache static background images at different angles during the initial calibration process. In particular, to reduce the number of cached images, assuming that the collection range of the target camera is a, first, the image range collected by the target camera at a fixed angle is b, then (a / b) static background images can be collected, where when the value of a / b is not an integer, it can be rounded up. Then, the (a / b) static background images are spliced to obtain an overall static background image for the collection range a. Next, according to the current camera angle corresponding to the target camera, the target static background image corresponding to the current camera angle is cropped from the overall background image.
[0044] Step S3: determining the image difference between each initial image in the initial image sequence and the target background image to generate a difference image, thereby obtaining a difference image sequence.
[0045] In practice, the difference image = initial image - target background image. In particular, since the initial image and the target background image are both captured by the target camera, the image specifications of the initial image and the target background image are consistent. For each pixel point in the initial image, the image difference is determined by the following formula:
[0046]
[0047] where i represents the horizontal coordinate number, j represents the vertical coordinate number, P s represents the pixel value corresponding to the pixel point in the initial image, P t represents the pixel value corresponding to the pixel point in the target background image, P d represents the pixel value corresponding to the pixel point in the difference image, represents the pixel value corresponding to the pixel point in the i-th row and j-th column of the initial image, represents the pixel value corresponding to the pixel point in the i-th row and j-th column of the target background image, represents the pixel value corresponding to the pixel point in the i-th row and j-th column of the difference image, represents the weight coefficient, the weight coefficient The value range of the weight coefficient The default value can be 0.9, and σ represents the noise threshold. Since the value range of the pixel value is [0, 255], the default value of the noise threshold σ can be 5.
[0048] Step S4: selecting a difference image that meets the screening condition from the difference image sequence as a candidate image, thereby obtaining a candidate image sequence.
[0049] The screening condition is that the number of pixel points containing non-zero pixel values in the difference image is greater than a preset number. Specifically, by setting the preset number, the initial image is preliminarily screened to determine whether it contains a living object. In particular, assuming that the image size of the initial image is HxW, the preset number can be set to (HxW) / 4.
[0050] Step S5: for each candidate image in the candidate image sequence, the following processing steps are performed:
[0051] Step S51: determining whether the candidate image contains a non-background object.
[0052] Wherein, since the purpose of step S51 is to determine whether the candidate image contains a person, face recognition, human body recognition, etc. can be used to determine whether the candidate image contains a non-background object. Specifically, face recognition can be performed on the candidate image using a FaceNet model. Human body recognition can also be performed on the candidate image using YOLO-v5 (You Only Look Once-version 5).
[0053] Step S52: In response to the candidate image containing a non-background object, object feature extraction is performed on the candidate image to obtain target object features.
[0054] In practice, for example, when face recognition is performed on the candidate image using the FaceNet model, a plurality of linearly connected fully connected layers can be set after the FaceNet model to map the face features extracted by the FaceNet model to a one-dimensional feature vector as the target object features. For another example, when human body recognition is performed on the candidate image using the Tiny-YOLO model, a plurality of linearly connected fully connected layers can be set after the Tiny-YOLO model to map the face features extracted by the Tiny-YOLO model to a one-dimensional feature vector as the target object features. In the present disclosure, the Tiny-YOLO model is preferred for constructing the target object features. The reason is that the face may not be included in the candidate image with the movement of the person, and the body occupies a higher proportion of the whole body, so the success rate of recognizing and extracting the target object features is higher.
[0055] Step S53: Perform object feature matching between the target object features and the object features corresponding to the to-be-identified object pre-stored corresponding to the to-be-identified object to obtain an object feature matching degree.
[0056] Wherein, the object features corresponding to the to-be-identified object can be pre-acquired and stored. In particular, the object features corresponding to the to-be-identified object are also one-dimensional feature vectors.
[0057] In practice, the feature similarity between the target object features and the object features pre-stored corresponding to the to-be-identified object can be determined by calculating the cosine similarity as the object feature matching degree.
[0058] Step S54: In response to the object feature matching degree being greater than a preset matching degree threshold, the image position of the candidate image in the initial image sequence is determined as the target position.
[0059] Optionally, the target camera can have at least one camera angle change in the target time period, and thus the initial image sequence can be divided into a plurality of initial image groups for a time interval of each camera angle change in the at least one camera angle change, each initial image group corresponds to a camera angle change, and the steps S2 to S5 are repeatedly executed.
[0060] In step 103, the initial image sequence is image extracted according to the target position sequence to obtain a target image sequence.
[0061] In some embodiments, the execution subject can image extract the initial image sequence according to the target position sequence to obtain a target image sequence.
[0062] The target image corresponds to the target position. Since the target position represents the sequence position of the initial image containing the to-be-recognized object in the initial image sequence, the initial image corresponding to the target position in the initial image sequence can be taken as the target image to obtain the target image sequence.
[0063] As an example, the following pseudo code can be specifically referred to:
[0064] Target_Img_List = []
[0065] for i in range(len(Target_Position_List)):
[0066] Target_Img_List.append(Initial_Img_List[Target_Position_List[i]])
[0067] Wherein, “Target_Img_List” represents the target image sequence, “Target_Position_List” represents the target position sequence, and “Initial_Img_List” represents the initial image sequence.
[0068] In step 104, the object description information set corresponding to the to-be-recognized object is generated according to the pre-trained object description information generation model and the target image sequence.
[0069] In some embodiments, the execution subject can image extract the initial image sequence according to the target position sequence to obtain a target image sequence.
[0070] The object description information includes a predicted action type, an identified action type, and emotion description information. The identified action type represents the action type of the to-be-identified object in the target image of the current frame. The predicted action type represents the action type of the to-be-identified object in the target image. The emotion description information represents the emotion change of the to-be-identified object. The object description information generation model is a machine learning model for generating object description information.
[0071] As an example, the target image sequence can include target image I1, target image I2, target image I3, target image I4, and target image I5. For target image I4, first, the object description information generation model can perform action type identification and emotion prediction on target image I4 to obtain the identified action type and the emotion description information included in the object description information corresponding to target image I4. Then, according to target image I1, target image I2, target image I3, and the object description information generation model, the predicted action type included in the object description information corresponding to target image I5 is generated.
[0072] Optionally, the object description information generation model includes an action type identification network, an action type prediction network, and an emotion identification network. The action type identification network is used to identify the action type of the to-be-identified object in a single target image. The action type prediction network is used to predict the action type of the to-be-identified object in the next frame of target image.
[0073] As an example, referring to Figure 2 The generation process of the object description information is shown in the figure. The target image sequence can include target image I1, target image I2, target image I3, target image I4, and target image I5. Taking target image I5 as an example, first, target image I5 is input into the action type identification network 203 included in the object description information generation model 201 to obtain the identified action type included in the object description information corresponding to target image I5. Then, target image I5 is input into the emotion identification network 204 to obtain the emotion description information included in the object description information corresponding to target image I5. Next, at least one target image (target image I1, target image I4, target image I3, target image I4) before target image I5 is input into the action type prediction network 202 to obtain the predicted action type included in the object description information corresponding to target image I5.
[0074] In some optional implementations of some embodiments, the execution subject generates a set of object description information corresponding to the to-be-identified object according to the pre-trained object description information generation model and the target image sequence, including:
[0075] For each target image in the target image sequence, the following identification steps are performed:
[0076] Step S1: performing action type recognition on the target image by an action type recognition network to obtain an identified action type included in the object description information corresponding to the target image.
[0077] The action type recognition network comprises a sub-action type recognition network A1, a sub-action type recognition network A2 and a sub-action type recognition network A3. The sub-action type recognition network A1, the sub-action type recognition network A2 and the sub-action type recognition network A3 all adopt the same “Encoder-Decoder” structure. The difference lies in that the Encoder part of the sub-action type recognition network A1 comprises five Transformer Blocks. The Encoder part of the sub-action type recognition network A2 comprises nine Transformer Blocks. The Encoder part of the sub-action type recognition network A3 comprises thirteen Transformer Blocks. The Transformer Block comprises an LN (Layer Normalization) layer, an MHSA (Multi Head Self Attention) layer, an LN layer and an FFN (Feed Forward Network) network. For the Transformer Block, the feature is input to the MHSA layer after passing through the first LN layer. Meanwhile, the output of the first LN layer and the output of the MHSA layer are superimposed as the input of the second LN layer. The output of the second LN layer is the input of the FFN network. Finally, the superimposed result of the output of the first LN layer and the output of the MHSA layer and the output of the FFN network are superimposed as the output of the Transformer Block. In particular, the five Transformer Blocks included in the Encoder part of the sub-action type recognition network A1 are consistent with the first five Transformer Blocks of the nine Transformer Blocks included in the Encoder part of the sub-action type recognition network A2. The nine Transformer Blocks included in the Encoder part of the sub-action type recognition network A2 are consistent with the first nine Transformer Blocks of the thirteen Transformer Blocks included in the Encoder part of the sub-action type recognition network A3. The Decoder part of the sub-action type recognition network A1, the Decoder part of the sub-action type recognition network A2 and the Decoder part of the sub-action type recognition network A3 all have a Deconv layer, a BN (Batch Normalization) layer, a ReLU activation function, a Deconv layer, a BN layer and a ReLU activation function.
[0078] Specifically, first, the target image is subjected to Patch Embedding processing. Then, the image block set after the Patch Embedding processing is input into the sub-action type recognition network A1, and an identified action type P1 is generated. When the confidence of the identified action type P1 is less than a confidence threshold, the features extracted by the Encoder part in the sub-action type recognition network A1 are taken as the input of the 6th Transformer Block in the sub-action type recognition network A2, and an identified action type P2 is generated through the sub-action type recognition network A2. When the confidence of the identified action type P2 is less than the confidence threshold, the features extracted by the Encoder part in the sub-action type recognition network A2 are taken as the input of the 10th Transformer Block in the sub-action type recognition network A2, and an identified action type P3 is generated through the sub-action type recognition network A3, and the identified action type P3 is taken as the identified action type corresponding to the target image. When the confidence of the identified action type P1 is greater than the confidence threshold, the identified action type P1 is taken as the identified action type corresponding to the target image, and no further action type recognition is performed through the sub-action type recognition network A2 and the sub-action type recognition network A3. When the confidence of the identified action type P2 is greater than the confidence threshold, the identified action type P2 is taken as the identified action type corresponding to the target image, and no further determination of the identified action type is performed through the sub-action type recognition network A3.
[0079] In practice, due to the different information feature abundances contained in the images, the conventional fixed network structure mode has the problem of feature redundancy processing, that is, a high-accuracy action type recognition can be achieved through several shallow feature processing layers, but the fixed network structure limits the need for redundant feature processing of the features. Therefore, the present disclosure determines the identified action type in sequence according to the structural complexity by setting three sub-action type recognition networks with different structural complexities. This mode can reduce the number of feature processing and improve the recognition speed. In addition, in order to avoid the problem of repeated feature extraction when the recognition confidence of the previous sub-action type recognition network is small and the next sub-action type recognition network needs to be recognized, the three sub-action type recognition network architectures have overlapping parts to ensure direct input of the features. Through this mode, the determination of the identified action type is realized at a low complexity and at a high speed.
[0080] Step S2: According to the above emotion recognition network, the object emotion of the target image is recognized to obtain the emotion label corresponding to the target image.
[0081] The emotion recognition network adopts a FaceNet network as a backbone network and connects an FC (Fully Connected Layer) layer as an emotion label classifier.
[0082] In practice, since the action type recognition network has extracted the key points contained in the face during the identification of the action type determination process, first, the region surrounded by the key points contained in the face can be taken as a region of interest. Then, the pixel values of the pixel points outside the region of interest in the target image are set to 0 to obtain an updated target image. Finally, the updated target image is input into the emotion recognition network to obtain the emotion label corresponding to the target image. In this way, the data processing amount can be reduced, and the face recognition process directly reuses the face region of interest located by the action type recognition network.
[0083] Step S3: In response to the first target image in the sequence of target images and non-target images, the predicted action type included in the object description information corresponding to the target image is generated according to the action type prediction network and at least one target image in the sequence of target images located before the target image.
[0084] The action type prediction network adopts a MoveNet as a backbone network.
[0085] In particular, the action type recognition network generates an image containing a skeletal anchor point during the identification of the action type determination process. Therefore, after processing each target image in the sequence of target images, a sequence of skeletal anchor point images is obtained, and therefore at least one skeletal anchor point image corresponding to at least one target image located before the target image can be input into the action type prediction network to obtain the predicted action type included in the object description information corresponding to the target image. This way avoids the re-determination of the skeletal anchor point, thereby reducing the data processing amount.
[0086] Step S4: In response to the first target image in the sequence of target images and non-target images, the emotion description information included in the object description information corresponding to the target image is generated according to the emotion label corresponding to at least one target image located before the target image in the sequence of target images and the emotion label corresponding to the target image.
[0087] The emotion description information can be stored in a sequence form to store the emotion labels corresponding to each target image before the target image.
[0088] Step S5: in response to the target image being the first target image in the target image sequence, determining the recognition action type corresponding to the target image as the predicted action type included in the object description information corresponding to the target image, and determining the current emotion label as the emotion description information included in the object description information corresponding to the target image.
[0089] Step 105: generating mild cognitive impairment evaluation information for the to-be-identified object according to the object description information set.
[0090] In some embodiments, the execution subject can generate mild cognitive impairment evaluation information for the to-be-identified object according to the object description information set.
[0091] When the to-be-identified object has mild cognitive impairment, it may have one or more cognitive function decline, action forgetfulness, abnormal emotional changes, and the like. Therefore, by judging the difference between the recognition action type and the predicted action type included in the object description information, the action forgetfulness of the to-be-identified object can be measured. At the same time, the emotional changes of the to-be-identified object can be linearly depicted according to the emotion label included in the object description information.
[0092] In practice, first, for the recognition action type and the predicted action type included in each object description information, the difference between the recognition action type and the predicted action type can be determined by calculating the similarity between them. The proportion of the object description information corresponding to the difference greater than the preset difference threshold in the object description information set is counted. Second, since the emotion label included in the object description information linearly depicts the emotional changes of the to-be-identified object, the quantification of the emotional changes of the to-be-identified object can be formed by counting the number of wave peaks and wave troughs. Finally, the mild cognitive impairment evaluation information for the to-be-identified object is formed through the decision tree model. Specifically, the mild cognitive impairment evaluation information can correspond to multiple evaluation levels.
[0093] Step 106: synchronizing the mild cognitive impairment evaluation information to the smart elderly care platform.
[0094] In some embodiments, the execution subject can synchronize the mild cognitive impairment evaluation information to the smart elderly care platform.
[0095] The smart elderly care platform can be a platform for real-time monitoring of user (elderly) physical monitoring levels.
[0096] In some optional implementations of some embodiments, the method further includes:
[0097] Step S1: determining whether there is a historical user portrait corresponding to the to-be-identified object.
[0098] The wisdom pension platform can include a user portrait module. The user portrait module can be used to store data representing the physical condition of the elderly in the form of a user portrait.
[0099] In practice, the identity of the to-be-identified object can be used to search in the user portrait module to determine whether there is a historical user portrait corresponding to the to-be-identified object.
[0100] Step S2: In response to the existence, updating the historical user portrait corresponding to the to-be-identified object according to the mild cognitive impairment evaluation information to obtain an updated user portrait.
[0101] In practice, the historical evaluation of mild cognitive impairment in the historical user portrait corresponding to the to-be-identified object can be updated according to the mild cognitive impairment evaluation information to obtain an updated user portrait.
[0102] Step S3: In response to the non-existence, generating an updated user portrait according to the mild cognitive impairment evaluation information and the basic object information corresponding to the to-be-identified object.
[0103] In practice, when there is no existence, the updated user portrait can be constructed in combination with the basic object information corresponding to the to-be-identified object and the mild cognitive impairment evaluation information. The basic object information can include object age, object gender, object disease history, object medical history, etc.
[0104] Step S4: Determining whether to initiate a mild cognitive impairment warning according to the updated user portrait and the pre-constructed warning matching rule.
[0105] The warning matching rule can be an automatic triggering rule set by a person for age, gender, disease history, and mild cognitive impairment evaluation information.
[0106] In practice, a corresponding rule trigger can be created in combination with the warning matching rule to automatically match the updated user portrait and the pre-constructed warning matching rule and trigger the mild cognitive impairment warning.
[0107] Step S5: In response to initiating the mild cognitive impairment warning, automatically initiating the mild cognitive impairment warning to the target terminal.
[0108] The target terminal includes a first target terminal, a second target terminal, and a third target terminal. The first target terminal is a terminal bound to the direct relative identity of the to-be-identified object. The second target terminal is a terminal bound to the care object identity of the to-be-identified object. The third target terminal is a terminal bound to the diagnosis and treatment object identity of the to-be-identified object. The care object can be an object that provides daily care to the to-be-identified object, such as a nurse. The diagnosis and treatment object can be an object that provides disease diagnosis and treatment to the to-be-identified object, such as a doctor.
[0109] Step S6: Recall the auxiliary diagnosis suggestion information corresponding to the updated user portrait to obtain an auxiliary diagnosis suggestion information set.
[0110] In practice, the diagnosis suggestions made by different doctors for different situations can be uniformly stored in the diagnosis suggestion pool. Therefore, the top K auxiliary diagnosis suggestion information matching the updated user portrait can be recalled from the diagnosis suggestion pool as the auxiliary diagnosis suggestion information set.
[0111] Step S7: Synchronize the auxiliary diagnosis suggestion information set to the third target terminal.
[0112] In practice, the auxiliary diagnosis suggestion information set is sent to the third target terminal to assist the diagnosis and treatment object in further determining the mild cognitive impairment of the to-be-identified object.
[0113] Step S8: Synchronize the target auxiliary diagnosis suggestion information to the second target terminal.
[0114] The target auxiliary diagnosis suggestion information is the auxiliary diagnosis suggestion information selected by the diagnosis and treatment object from the auxiliary diagnosis suggestion information set and matching the to-be-identified object.
[0115] In practice, the target auxiliary diagnosis suggestion information is sent to the second target terminal, so that the care object can initiate the mild cognitive impairment determination and diagnosis process for the to-be-identified object according to the target auxiliary diagnosis suggestion information.
[0116] The above various embodiments of the present disclosure have the following beneficial effects: through the image recognition-based mild cognitive impairment identification method of some embodiments of the present disclosure, effective identification of changes in the patient's mild cognitive impairment is achieved. Specifically, the reason that cannot be identified in time is that the onset process of mild cognitive impairment has the characteristics of long periodicity, and the regular return visit diagnosis by using the conventional medical diagnosis method has poor recognition effect on mild cognitive impairment in the window period, which makes it difficult to discover the patient's mild cognitive impairment changes in time, thereby increasing the risk of developing into Alzheimer's disease and the like. For example, assuming that the patient needs to be diagnosed at T1 time point and T2 time point, when the patient has a mild cognitive impairment change in the time window formed by T1 time point and T2 time point, it is difficult to discover in time, especially as the window period increases, the risk of deterioration further increases. Based on this, the image recognition-based mild cognitive impairment identification method of some embodiments of the present disclosure first acquires an initial image sequence for a to-be-identified object, wherein the to-be-identified object is an object corresponding to an age greater than a preset age, which is pre-marked to be identified for mild cognitive impairment. Secondly, object recognition is performed on the initial images in the initial image sequence to obtain a target position sequence, wherein the target position represents the sequence position of the initial image containing the to-be-identified object in the initial image sequence. Then, according to the target position sequence, image extraction is performed on the initial image sequence to obtain a target image sequence, wherein the target image corresponds to the target position. In practice, the to-be-identified object is the subject of mild cognitive impairment identification, and it does not remain stationary, so there may be cases where the to-be-identified object is not included in the picture. Since the onset process of mild cognitive impairment has the characteristics of long periodicity, a large number of images need to be identified and processed. In order to avoid unnecessary consumption of computing resources, the target position is determined by object recognition, which filters out invalid image frames. Further, according to the pre-trained object description information generation model and the target image sequence, an object description information set corresponding to the to-be-identified object is generated, wherein the object description information includes: predicted action type, identified action type and emotion description information, wherein the emotion description information represents the emotional changes of the to-be-identified object. In practice, since the patient with mild cognitive impairment has one or more cognitive function declines, the present disclosure generates object description information (set) for the to-be-identified object changing over time from two dimensions of behavior and emotion. Then, according to the object description information set, mild cognitive impairment evaluation information for the to-be-identified object is generated. In this way, the mild cognitive impairment evaluation for the to-be-identified object is automatically generated from two dimensions of behavior and emotion. Finally, the mild cognitive impairment evaluation information is synchronized to the smart elderly care platform. In this way, effective identification of changes in the patient's mild cognitive impairment is achieved.
[0117] Further referenceFigure 3 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an image recognition-based mild cognitive impairment recognition device for the elderly, which device embodiments correspond to those method embodiments, and the image recognition-based mild cognitive impairment recognition device for the elderly can be specifically applied in various electronic devices. Figure 1
[0118] As shown in Figure 3 , the image recognition-based mild cognitive impairment recognition device for the elderly 300 of some embodiments includes an acquisition unit 301, an object recognition unit 303, an image extraction unit 303, a first generation unit 304, a second generation unit 305, and a synchronization unit 306. Among them,
[0119] The acquisition unit 301 is configured to acquire an initial image sequence for a to-be-recognized object, wherein the to-be-recognized object is an object corresponding to an age greater than a preset age, which is pre-labeled to be subjected to mild cognitive impairment recognition; the object recognition unit 302 is configured to perform object recognition on the initial images in the initial image sequence to obtain a target position sequence, wherein the target position represents the sequence position of the initial image containing the to-be-recognized object in the initial image sequence; the image extraction unit 303 is configured to perform image extraction on the initial image sequence according to the target position sequence to obtain a target image sequence, wherein the target image corresponds to the target position; the first generation unit 304 is configured to generate an object description information set corresponding to the to-be-recognized object according to a pre-trained object description information generation model and the target image sequence, wherein the object description information includes a predicted action type, a recognized action type, and emotion description information, wherein the emotion description information represents the emotional changes of the to-be-recognized object; the second generation unit 305 is configured to generate mild cognitive impairment evaluation information for the to-be-recognized object according to the object description information set; and the synchronization unit 306 is configured to synchronize the mild cognitive impairment evaluation information to the smart elderly care platform.
[0120] It can be understood that the units described in the image recognition-based mild cognitive impairment recognition device for the elderly 300 correspond to the respective steps in the method described with reference to Figure 1 . Therefore, the operations, features, and beneficial effects described above for the method also apply to the image recognition-based mild cognitive impairment recognition device for the elderly 300 and the units contained therein, which will not be described here.
[0121] Reference is made below to Figure 4 , which shows a structural schematic diagram of an electronic device (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the function and use range of the embodiments of the present disclosure. As shown in Figure 4 As shown in the structure shown in the embodiment, the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to perform any of the above methods. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the computer program in the non-volatile storage medium to run, which, when executed by the processor, can cause the processor to perform any of the above methods. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4 The structure shown in the embodiment is merely a block diagram of part of the structure related to the present disclosure scheme, and does not constitute a limitation on the computer device to which the present disclosure scheme is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0122] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0123] In one embodiment, the processor is configured to run a computer program stored in the memory to implement the following steps: obtaining an initial image sequence of a to-be-identified object, wherein the to-be-identified object is an object corresponding to an age greater than a preset age and pre-labeled as to be identified for mild cognitive impairment; performing object identification on an initial image in the initial image sequence to obtain a target position sequence, wherein the target position represents a sequence position of the initial image containing the to-be-identified object in the initial image sequence; performing image extraction on the initial image sequence according to the target position sequence to obtain a target image sequence, wherein the target image corresponds to the target position; generating an object description information set corresponding to the to-be-identified object according to a pre-trained object description information generation model and the target image sequence, wherein the object description information includes a predicted action type, an identified action type, and emotion description information, wherein the emotion description information represents an emotional change of the to-be-identified object; generating mild cognitive impairment evaluation information for the to-be-identified object according to the object description information set; and synchronizing the mild cognitive impairment evaluation information to the smart elderly care platform.
[0124] The embodiments of the present disclosure further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program includes program instructions. The method implemented by the program instructions can refer to the embodiments of the method of the present disclosure.
[0125] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like.
[0126] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article, or system that includes a list of elements not only includes those elements, but also includes other elements not explicitly listed, or inherent to such a process, method, article, or system. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article, or system that includes the element.
[0127] The above description is merely exemplary of some of the many possible embodiments of the present disclosure and of the principles thereof. It is to be understood that those skilled in the art will be able to devise various embodiments of the present disclosure without departing from the scope of the present disclosure as disclosed in the above description and attached claims, and that the scope of the present disclosure is not limited to the specific technical features described above. For example, the technical features described above can be replaced with other technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form other technical solutions.
Claims
1. A method for identifying mild cognitive impairment in the elderly based on image recognition, applied to a smart elderly care platform, characterized in that, include: Acquire an initial image sequence for the object to be identified, wherein the object to be identified is an object whose age is greater than a preset age and which is pre-labeled as an object to be identified for mild cognitive impairment; Object recognition is performed on the initial images in the initial image sequence to obtain a target position sequence, wherein the target position represents the sequence position of the initial image containing the object to be identified in the initial image sequence; Based on the target location sequence, image extraction is performed on the initial image sequence to obtain a target image sequence, wherein the target image corresponds to the target location; Based on the pre-trained object description information generation model and the target image sequence, an object description information set corresponding to the object to be identified is generated. The object description information includes: predicted action type, identified action type and emotion description information, wherein the emotion description information represents the emotion changes of the object to be identified. Based on the object description information set, mild cognitive impairment assessment information is generated for the object to be identified; The mild cognitive impairment assessment information will be synchronized to the smart elderly care platform.
2. The method according to claim 1, characterized in that, The method further includes: Determine whether a historical user profile exists corresponding to the object to be identified; In response to the existence, the user profile of the historical user profile corresponding to the object to be identified is updated based on the mild cognitive impairment assessment information to obtain the updated user profile; In response to the absence of the object, an updated user profile is generated based on the mild cognitive impairment assessment information and the basic object information corresponding to the object to be identified. Based on the updated user profile and the pre-built warning matching rules, determine whether to initiate a mild cognitive impairment warning; In response to initiating a mild cognitive impairment warning, an automatic mild cognitive impairment warning is sent to target terminals, including: a first target terminal, a second target terminal, and a third target terminal. The first target terminal is a terminal bound to the identity of the immediate family member of the person to be identified, the second target terminal is a terminal bound to the identity of the caregiver of the person to be identified, and the third target terminal is a terminal bound to the identity of the patient receiving medical treatment for the person to be identified.
3. The method according to claim 2, characterized in that, The method further includes: Recall the auxiliary diagnostic suggestion information corresponding to the updated user profile to obtain a set of auxiliary diagnostic suggestion information; Synchronize the set of auxiliary diagnostic suggestion information to the third target terminal; The target auxiliary diagnostic suggestion information is synchronized to the second target terminal, wherein the target auxiliary diagnostic suggestion information is the auxiliary diagnostic suggestion information selected by the treatment subject from the auxiliary diagnostic suggestion information set and matched with the object to be identified.
4. The method according to claim 3, characterized in that, The step of performing object recognition on the initial images in the initial image sequence to obtain the target location sequence includes: Determine whether there is a change in camera angle of the target camera within a target time period, wherein the target camera is the camera that acquires the initial image sequence, and the target time period is the acquisition time period corresponding to the initial image sequence; In response to the absence of camera angle change, a target background image is acquired, wherein the target background image is a static background image at the current camera angle corresponding to the target camera; The image difference between each initial image in the initial image sequence and the target background image is determined to generate a difference image, thus obtaining a difference image sequence; From the difference image sequence, difference images that meet the filtering criteria are selected as candidate images to obtain a candidate image sequence. The filtering criteria are: the number of pixels with non-zero pixel values contained in the difference image is greater than a preset number. For each candidate image in the candidate image sequence, perform the following processing steps: Determine whether the candidate image contains a non-background object; In response to the presence of a non-background object in the candidate image, object features are extracted from the candidate image to obtain the target object features; The target object features are matched with the pre-stored object features corresponding to the object to be identified to obtain the object feature matching degree. In response to an object feature matching degree greater than a preset matching degree threshold, the image position of the candidate image in the initial image sequence is determined as the target position.
5. The method according to claim 4, characterized in that, The object description information generation model includes: an action type recognition network, an action type prediction network, and an emotion recognition network. The action type recognition network is used to identify the action type of the object to be identified within a single target image. The action type prediction network is used to predict the action type of the object to be identified in the next frame of the target image. The step of generating an object description information set corresponding to the object to be identified based on the pre-trained object description information generation model and the target image sequence includes: For each target image in the target image sequence, perform the following recognition steps: By using an action type recognition network, the target image is subjected to action type recognition to obtain the object description information corresponding to the target image, including the recognized action type. Based on the emotion recognition network, object emotion recognition is performed on the target image to obtain the emotion label corresponding to the target image; In response to the first target image in the target image-non-target image sequence, the predicted action type is generated based on the action type prediction network and at least one target image in the target image sequence that precedes the target image; In response to the first target image in the target image-non-target image sequence, based on the action type prediction network and the emotion label corresponding to at least one target image in the target image sequence that is preceding the target image and the emotion label corresponding to the target image, the object description information corresponding to the target image is generated, including the emotion description information. In response to the target image being the first target image in the target image sequence, the recognition action type corresponding to the target image is determined as the predicted action type included in the object description information corresponding to the target image, and the current emotion tag is determined as the emotion description information included in the object description information corresponding to the target image.
6. A device for identifying mild cognitive impairment in the elderly based on image recognition, characterized in that, include: The acquisition unit is configured to acquire an initial image sequence for an object to be identified, wherein the object to be identified is an object that is pre-labeled as being to be identified for mild cognitive impairment and whose corresponding age is greater than a preset age; An object recognition unit is configured to perform object recognition on the initial images in the initial image sequence to obtain a target position sequence, wherein the target position represents the sequence position of the initial image containing the object to be recognized in the initial image sequence. An image extraction unit is configured to extract images from an initial image sequence based on the target location sequence to obtain a target image sequence, wherein the target images correspond to target locations; The first generation unit is configured to generate a set of object description information corresponding to the object to be identified based on a pre-trained object description information generation model and the target image sequence. The object description information includes: predicted action type, identified action type and emotion description information, wherein the emotion description information represents the emotion change of the object to be identified. The second generation unit is configured to generate mild cognitive impairment assessment information for the object to be identified based on the object description information set. The synchronization unit is configured to synchronize the mild cognitive impairment assessment information to the smart elderly care platform.
7. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.
8. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Face recognition, non-inductive heart rate recognition, running posture and expression recognition integrated method and system
CN113869160A
Skeleton sequence-based old person behavior identification method in infrared video
CN114724251A
Old people cognitive disorder detection system and method based on behavior analysis
CN117158904A
Monitoring image processing method and system based on machine vision
CN118154635A
Multi-dimensional assessment method and system for disability risk of old people
CN118711825A