Artificial intelligence-based personnel missing judgment method, device and equipment and medium

A multi-modal neural network model for elderly wandering detection improves accuracy and efficiency by integrating location, image, and audio data to enhance safety and data reliability.

CN120316705APending Publication Date: 2025-07-15PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510374462.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing technology cannot accurately determine whether the elderly or patients are missing, resulting in large errors in medical data and the inability to deal with the risk of the elderly being lost in a timely manner.

Method used

By obtaining the position data, image data and voice data of the surveillance personnel, using the multimodal neural network model for feature extraction, fusion and encoding processing, outputting the loss judgment results, and improving the discrimination accuracy.

Benefits of technology

It improves the efficiency and accuracy of judging the loss of the elderly or patients, reduces medical data errors, and ensures the safety of the elderly and the accuracy of data statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316705A_ABST
    Figure CN120316705A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence-based person missing discrimination method and device, equipment and a medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining first position data, first image data and first voice data of an environment where a monitored person is located; performing feature extraction on the first position data, the first image data and the first voice data through a feature extraction network in a preset person missing discrimination model to obtain a position feature, an image feature and a voice feature; performing fusion processing on the position features, the image features and the voice features through a feature fusion network to obtain fusion features; performing encoding processing on the fusion feature through an encoder to obtain a first feature vector; and outputting a lost judgment result of the monitored person based on the first feature vector through a judgment layer. According to the invention, the efficiency and accuracy of lost judgment are greatly improved. The method can be applied to the fields of medical health, old-age care and the like, and the safety of the monitored person is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and medium for judging the loss of people based on artificial intelligence. Background Art

[0002] With the further aggravation of the trend of population aging, the problem of the care of the elderly has become an issue that everyone must face along with the development of the city. As the elderly grow older, they will more or less suffer from some geriatric diseases (such as memory decline and Alzheimer's disease, etc.), resulting in the risk of getting lost during daily activities such as buying groceries and taking walks, and being unable to return to their place of residence. In addition, in the scenario where the elderly live in a nursing home, when the elderly arrive in a new environment, they are only familiar with the environment near the nursing home. When the elderly go a little further away, they will have an accident of getting lost. In the field of medical and health, in medical data statistics, it is necessary to count the number of times whether the patient gets lost in daily life, but there is currently no good way to accurately determine whether the patient gets lost, resulting in large errors in medical data.

[0003] Therefore, how to improve the accuracy of the judgment result of the loss of the monitored person is an urgent problem to be solved at present. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, equipment and medium for judging the loss of people based on artificial intelligence, aiming to improve the accuracy of the judgment result of the loss of the monitored person.

[0005] In a first aspect, this application provides a method for judging the loss of people based on artificial intelligence, and the method for judging the loss of people includes the following steps:

[0006] Obtain first data to be detected in the environment where the monitored person is located, and the first data to be detected includes first position data, first image data and first voice data;

[0007] Respectively perform feature extraction on the first position data, the first image data and the first voice data through a feature extraction network in a preset person loss judgment model, and obtain position features, image features and voice features. The person loss judgment model is obtained by pre-training a multi-modal neural network model based on multiple training samples, and the training samples include sample position data, sample image data, sample voice data and labeled loss judgment results;

[0008] Perform fusion processing on the position features, the image features and the voice features through a feature fusion network in the person loss judgment model to obtain fusion features;

[0009] Encode the fusion feature through the encoder in the person missing discrimination model to obtain a first feature vector;

[0010] Output the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model.

[0011] In a second aspect, the present application further provides a person missing discrimination device, which includes an acquisition module, a feature extraction module, a fusion processing module, an encoding processing module, and an output module, where:

[0012] The acquisition module is configured to acquire first detection data of the environment where the person under guardianship is located, and the first detection data includes first position data, first image data, and first voice data;

[0013] The feature extraction module is configured to respectively extract features from the first position data, the first image data, and the first voice data through a feature extraction network in a preset person missing discrimination model, to obtain position features, image features, and voice features. The person missing discrimination model is pre-trained based on multiple training samples for a multi-modal neural network model, and the training samples include sample position data, sample image data, sample voice data, and labeled missing determination results;

[0014] The fusion processing module is configured to perform fusion processing on the position features, the image features, and the voice features through a feature fusion network in the person missing discrimination model to obtain a fusion feature;

[0015] The encoding processing module is configured to encode the fusion feature through the encoder in the person missing discrimination model to obtain a first feature vector;

[0016] The output module is configured to output the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model.

[0017] In a third aspect, the present application further provides a computer device, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the person missing discrimination method as described above are implemented.

[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the person missing discrimination method as described above are implemented.

[0019] The present application provides a method, apparatus, device and medium for judging the loss of a person based on artificial intelligence. In the present application, first detection data of the environment where the person under guardianship is located is obtained, and the first detection data includes first position data, first image data and first voice data; the feature extraction network in the preset person loss judgment model is used to respectively extract features from the first position data, the first image data and the first voice data to obtain position features, image features and voice features. The person loss judgment model is obtained by pre-training a multi-modal neural network model based on a plurality of training samples, and the training samples include sample position data, sample image data, sample voice data and labeled loss judgment results; the feature fusion network in the person loss judgment model is used to perform fusion processing on the position features, the image features and the voice features to obtain fusion features; the encoder in the person loss judgment model is used to perform encoding processing on the fusion features to obtain a first feature vector; the judgment layer in the person loss judgment model outputs the loss judgment result of the person under guardianship based on the first feature vector. In the present application, by identifying the first position data, the first image data and the first voice data of the environment where the person under guardianship is located, the environmental perception ability can be improved, and then the efficiency and accuracy of loss judgment can be improved. Then, based on the multi-modal neural network model, loss judgment is performed on the first position data, the first image data and the first voice data. The multi-modal neural network model can improve the accuracy of data processing, and thus greatly improve the efficiency and accuracy of loss judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flow chart of a method for judging the loss of a person provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic flow chart of another method for judging the loss of a person provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of a person loss judgment model provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic block diagram of a person loss judgment device provided by an embodiment of the present application;

[0025] Figure 5Schematic block diagram of another device for discriminating the loss of a person provided in an embodiment of the present application;

[0026] Figure 6 Schematic block diagram of the structure of a computer device provided in an embodiment of the present application.

[0027] The realization of the purpose, functional characteristics and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0029] The flowcharts shown in the accompanying drawings are only illustrative, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0030] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, sense the environment, acquire knowledge and use the knowledge to obtain the best results of theory, method, technology and application system.

[0031] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0032] With the further intensification of the trend of population aging, the problem of elderly care has become an issue that everyone must face along with the development of the city. As the elderly grow older, they will more or less suffer from some geriatric diseases (such as memory decline and Alzheimer's disease, etc.), resulting in the risk of getting lost during daily activities such as grocery shopping and taking walks and being unable to return to their place of residence. Moreover, in the scenario where the elderly live in a nursing home, when they arrive in a new environment, they are only familiar with the environment near the nursing home. When the elderly walk a little further away, they will have an accident of getting lost. In the field of medical health, in medical data statistics, it is necessary to count the number of times whether a patient gets lost in daily life, but there is currently no good way to accurately determine whether a patient gets lost, resulting in a large error in medical data.

[0033] To solve the above problems, an embodiment of the present application provides a method, device, equipment and medium for discriminating the loss of a person based on artificial intelligence. Among them, the method for discriminating the loss of a person based on artificial intelligence includes: obtaining first detection data of the environment where the person under guardianship is located, and the first detection data includes first position data, first image data and first voice data;

[0034] Respectively performing feature extraction on the first position data, the first image data and the first voice data through a feature extraction network in a preset person loss discrimination model to obtain position features, image features and voice features. The person loss discrimination model is pre-trained on a multi-modal neural network model based on multiple training samples. The training samples include sample position data, sample image data, sample voice data and labeled loss determination results; performing fusion processing on the position features, the image features and the voice features through a feature fusion network in the person loss discrimination model to obtain fusion features; performing encoding processing on the fusion features through an encoder in the person loss discrimination model to obtain a first feature vector; and outputting a loss determination result of the person under guardianship based on the first feature vector through a determination layer in the person loss discrimination model.

[0035] Among them, the method for discriminating the loss of a person can be applied to a computer device, and the computer device can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant and a wearable device, etc.

[0036] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of a method for discriminating the loss of a person provided by an embodiment of the present application.

[0038] As shown Figure 1 in and

[0039] , the method for determining the loss of a person includes steps S101 to S105.

[0039] Step S101: Obtain first data to be detected in the environment where the person under guardianship is located. The first data to be detected includes first position data, first image data, and first voice data.

[0040] Among them, the first data to be detected is the position data, image data, and voice data of the environment where the person under guardianship is located. The position data is the positioning of the person under guardianship, the image data is the image of the environment where the person under guardianship is located, and the voice data is the voice of the person under guardianship and / or the voice in the surrounding environment.

[0041] In some embodiments, the person under guardianship wears a smart wearable device, and the first position data, first image data, and first voice data in the environment where the person under guardianship is located are obtained through the smart wearable device. Among them, the smart wearable device can be selected according to the actual situation, and the embodiments of the present application do not make specific limitations in this regard. For example, the smart wearable device can be an AR glasses and a smart watch, etc. Through the smart wearable device, the first position data, first image data, and first voice data of the person under guardianship can be accurately obtained, greatly improving the efficiency and accuracy of guardianship of the person under guardianship.

[0042] Exemplarily, the person under guardianship wears AR glasses. The position in the environment where the person under guardianship is located is obtained through the positioning system of the AR glasses to obtain the first position data; the image of the environment where the person under guardianship is located is obtained through the camera of the AR glasses to obtain the first image data; and the sound in the environment where the person under guardianship is located is obtained through the sound collection system of the AR glasses to obtain the first voice data.

[0043] In some embodiments, the position data in the environment where the person under guardianship is located is obtained through the smart wearable device, and the position data is preprocessed to obtain the first position data. By preprocessing the position data, the efficiency and accuracy of subsequent loss determination can be improved.

[0044] It should be noted that the methods for preprocessing the position data include but are not limited to noise filtering and position conversion, etc. The noise filtering can be Kalman filtering or particle filtering, and the position conversion can be the conversion between geographical coordinates and local coordinate systems. The noise filtering and position conversion can be selected according to the actual situation, and the embodiments of the present application do not make specific limitations in this regard.

[0045] In some embodiments, the image data in the environment where the person under guardianship is located is obtained through the smart wearable device, and the image data is preprocessed to obtain the first image data. By preprocessing the image data, the efficiency and accuracy of subsequent loss determination can be improved.

[0046] It should be noted that the methods for preprocessing the image data include, but are not limited to, image size scaling, image contrast adjustment, image resolution adjustment, and so on.

[0047] In some embodiments, voice data of the environment where the monitored person is located is acquired by the smart wearable device, and the voice data is preprocessed to obtain first voice data. By preprocessing the voice data, the efficiency and accuracy of subsequent missing person determination can be improved.

[0048] It should be noted that the methods for preprocessing the voice data include, but are not limited to, noise reduction.

[0049] Step S102: Feature extraction networks in a preset missing person discrimination model are respectively used to extract features from the first position data, the first image data, and the first voice data, so as to obtain position features, image features, and voice features.

[0050] Among them, the missing person discrimination model is obtained by pre-training a multi-modal neural network model based on multiple training samples, and the training samples include sample position data, sample image data, sample voice data, and labeled missing person determination results.

[0051] In some embodiments, as Figure 2 shown, the method further includes steps S201 to S205.

[0052] Step S201: Obtain a sample data set, where the sample data set includes multiple sample data, and the sample data includes sample position data, sample image data, sample voice data, and labeled missing person determination results.

[0053] Among them, the sample data set includes multiple sample data, the sample data includes sample position data, sample image data, sample voice data, and labeled missing person determination results, and the labeled missing person determination results are labeled by users.

[0054] In some embodiments, the monitored person wears a smart wearable device, and sample position data, sample image data, and sample voice data of the environment where the monitored person is located are acquired by the smart wearable device. When the caregiver determines whether the monitored person is missing when the monitored person is at the sample position, a labeled missing person determination result is obtained. Among them, the smart wearable device can be selected according to actual situations, and the embodiments of the present application do not make specific limitations thereon. For example, the smart wearable device can be an AR glasses, a smart watch, and so on. By using the smart wearable device, the sample position data, sample image data, and sample voice data of the monitored person can be accurately acquired, greatly improving the accuracy of the missing person discrimination model.

[0055] In some embodiments, the intelligent wearable device is used to obtain the location data of the environment where the person under guardianship is located, and the location data is preprocessed to obtain sample location data. By preprocessing the location data, the accuracy of the trained personnel missing discriminant model can be improved.

[0056] It should be noted that the methods for preprocessing the location data include but are not limited to noise filtering and location conversion, etc. The noise filtering can be Kalman filtering or particle filtering, and the location conversion can be the conversion between geographic coordinates and the local coordinate system. The noise filtering and location conversion can be selected according to the actual situation, and the embodiments of the present application do not make specific limitations on this.

[0057] In some embodiments, the intelligent wearable device is used to obtain the image data of the environment where the person under guardianship is located, and the image data is preprocessed to obtain sample image data. By preprocessing the image data, the accuracy of the trained personnel missing discriminant model can be improved.

[0058] It should be noted that the methods for preprocessing the image data include but are not limited to image size scaling, image contrast adjustment, image resolution adjustment, etc.

[0059] In some embodiments, the intelligent wearable device is used to obtain the voice data of the environment where the person under guardianship is located, and the voice data is preprocessed to obtain sample voice data. By preprocessing the voice data, the accuracy of the trained personnel missing discriminant model can be improved.

[0060] In some embodiments, a scheme is executed to obtain the sample location data, sample image data, and sample voice data of the environment where the person under guardianship is located through the intelligent wearable device once. When the guardian determines whether the person under guardianship is missing when the person under guardianship is at the sample location, and obtains the labeled missing judgment result, a sample data can be obtained. Repeating this scheme can obtain a sample data set.

[0061] Step S202: Obtain a preset multi-modal neural network model, and select a sample data from the sample data set as the target sample data.

[0062] Among them, as Figure 3 shown, the multi-modal neural network model includes a feature extraction network, a feature fusion network, an encoder, and a determination layer. The feature extraction network is used to extract features from location data, image data, and voice data to obtain location features, image features, and voice features; the feature fusion network is used to perform fusion processing on the location features, image features, and voice features to obtain fusion features; the encoder is used to perform encoding processing on the fusion features to obtain feature vectors; the determination layer outputs a missing determination result based on the feature vectors.

[0063] It should be noted that the feature extraction network can be the Bert, ResNet50, and Wav2vec models, and the feature fusion network is Uni-MoE.

[0064] In some embodiments, a sample data is randomly selected from the sample dataset as the target sample data. Among them, the target sample data includes target sample location data, target sample image data, target sample voice data, and the labeled lost determination result. By selecting a sample data from the sample dataset, the target sample data can be accurately obtained.

[0065] Step S203: Input the sample location data, sample image data, and sample voice data in the target sample data into the preset multi-modal neural network model for training to obtain the predicted lost determination result.

[0066] In some embodiments, the feature extraction network is used to extract features from the sample location data, sample image data, and sample voice data respectively to obtain the first location feature, the first image feature, and the first voice feature. By using the feature extraction network to extract features from the sample location data, sample image data, and sample voice data respectively, the first location feature, the first image feature, and the first voice feature can be accurately obtained.

[0067] In some embodiments, the feature fusion network performs fusion processing on the first location feature, the first image feature, and the first voice feature to obtain the first fusion feature; the encoder performs encoding processing on the first fusion feature to obtain the third feature vector; the determination layer outputs the predicted lost determination result of the person under guardianship based on the third feature vector.

[0068] In some embodiments, before the determination layer outputs the predicted lost determination result of the person under guardianship based on the third feature vector, it further includes: obtaining a plurality of historical feature vectors, which are determined according to the second target sample data, and the acquisition time of the second target sample data is before the target sample; determining the fourth feature vector according to the third feature vector and the plurality of historical feature vectors; the determination layer outputs the predicted lost determination result of the person under guardianship based on the third feature vector, including: the determination layer outputs the predicted lost determination result based on the fourth feature vector.

[0069] In some embodiments, the method for determining the fourth feature vector based on the third feature vector and multiple historical feature vectors may be as follows: determine the similarity between the third feature vector and each historical feature vector; construct a similarity matrix based on the similarity between the third feature vector and each historical feature vector; encode the similarity matrix through an encoder to obtain the fourth feature vector. By constructing a similarity matrix for the third feature vector and multiple historical feature vectors and then encoding the similarity matrix through an encoder, the fourth feature vector can be accurately obtained.

[0070] In some embodiments, the method for determining the similarity between the third feature vector and each historical feature vector may be as follows: obtain a preset similarity formula, and the preset similarity formula is The s i,j is the similarity, the h i,k is the third feature vector, and the h j,k is the historical feature vector. Calculate the third feature vector and the historical feature vector based on the preset similarity formula to obtain a similarity; and calculate the third feature vector and each historical feature vector based on the preset similarity formula respectively to obtain multiple similarities.

[0071] In some embodiments, the method for constructing a similarity matrix based on the similarity between the third feature vector and each historical feature vector may be as follows: arrange the similarities according to the time sequence of the time points when the historical feature vectors corresponding to the similarities are generated to obtain the similarity matrix. By arranging the similarities according to the time sequence of the time points when the historical feature vectors are generated, the similarity matrix can be accurately obtained.

[0072] In some embodiments, the method for the determination layer to output a predicted wandering determination result based on the fourth feature vector may be as follows: perform feature calculation on the fourth feature vector through a preset prediction classification output formula to obtain a predicted classification output feature value; calculate the predicted classification output feature value through a preset wandering evaluation formula to obtain a predicted wandering evaluation score; if the predicted wandering evaluation score is greater than or equal to a preset score, determine that the wandering determination result is that the person under guardianship has wandered. Among them, the preset score can be set according to the actual situation, and the embodiments of the present application do not make specific limitations on this. For example, the preset score can be set to 60%. By processing the fourth feature vector by the determination layer, the predicted wandering determination result can be accurately obtained.

[0073] Exemplarily, the preset wandering evaluation formula is The E is the classification output feature value, and f k is the classification output function. Among them, in the present application, the K classification is used as an example for illustration. Of course, the f kIt can also be other types of functions. Based on the preset formula for evaluating the risk of getting lost, the predicted classification output feature values are calculated to obtain the predicted classification output feature values.

[0074] Exemplarily, the preset formula for evaluating the risk of getting lost is where the y i is the predicted score for evaluating the risk of getting lost, the a i is the classification output feature value, and the a j is the classification output feature value. It can be understood that the denominator of the preset formula for evaluating the risk of getting lost is the sum of multiple classification output feature values. Based on the preset formula for evaluating the risk of getting lost, the classification output feature values are calculated to obtain the score for evaluating the risk of getting lost.

[0075] Step S204: According to the predicted result of getting lost and the labeled result of getting lost, determine whether the preset multi-modal neural network model converges.

[0076] According to the predicted result of getting lost and the labeled result of getting lost, determine the loss value of the preset multi-modal neural network model. If the loss value is less than or equal to the preset loss value, it is determined that the preset multi-modal neural network model has converged; if the loss value is greater than the preset loss value, it is determined that the preset multi-modal neural network model has not converged. Among them, the preset loss value can be set according to the actual situation, and the embodiments of the present application do not make specific limitations on this. For example, the preset loss value can be set to 0.002.

[0077] In some embodiments, the method for determining the loss value of the preset multi-modal neural network model according to the predicted result of getting lost and the labeled result of getting lost may be: if the predicted result of getting lost is the same as the labeled result of getting lost, the current loss value is labeled as 0; if the predicted result of getting lost is different from the labeled result of getting lost, the current loss value is labeled as 1; obtain the historical loss value, which is the loss value calculated during the previous model training process; perform an addition operation on the current loss value and the historical loss value and divide the result by 2 to obtain the loss value.

[0078] Step S205: If the preset multi-modal neural network model has not converged, continue to execute the step of selecting a sample data from the sample dataset as the target sample data until the preset multi-modal neural network model converges to obtain a model for discriminating whether a person gets lost.

[0079] If the loss value is greater than the preset loss value, it is determined that the preset multi-modal neural network model has not converged, and continue to execute the step of selecting a sample data from the sample dataset as the target sample data until the loss value is less than or equal to the preset loss value, and it is determined that the preset multi-modal neural network model has converged to obtain a model for discriminating whether a person gets lost.

[0080] In some embodiments, the feature extraction network is used to extract features from the first location data, the first image data, and the first voice data respectively, to obtain location features, image features, and voice features. By using the feature extraction network to extract features from the first location data, the first image data, and the first voice data respectively, the location features, image features, and voice features can be accurately obtained.

[0081] It should be noted that the feature extraction network may include a Bert model, a ResNet50 model, and a Wav2vec model. The Bert model is used for extracting location features, the ResNet50 model is used for extracting image features, and the Wav2vec model is used for extracting voice features.

[0082] Step S103: The feature fusion network in the missing person discrimination model is used to perform fusion processing on the location features, the image features, and the voice features to obtain fusion features.

[0083] Among them, the feature fusion network may be a Uni-MoE model.

[0084] In some embodiments, the feature fusion network in the missing person discrimination model is used to perform fusion processing on the location features, the image features, and the voice features to obtain fusion features. By performing fusion processing on the location features, the image features, and the voice features, the fusion features can be accurately obtained, and the efficiency and accuracy of missing person determination can be effectively improved.

[0085] Step S104: The encoder in the missing person discrimination model is used to perform encoding processing on the fusion features to obtain a first feature vector.

[0086] The encoder in the missing person discrimination model is used to perform encoding processing on the fusion features to obtain a first feature vector. By using the encoder to perform encoding processing on the fusion features, the first feature vector can be accurately obtained.

[0087] Step S105: The determination layer in the missing person discrimination model outputs the missing determination result of the person under guardianship based on the first feature vector.

[0088] In some embodiments, multiple historical feature vectors are obtained. The historical feature vectors are determined according to the second data to be detected, and the acquisition time of the second data to be detected is before that of the first data to be detected. According to the first feature vector and the multiple historical feature vectors, a second feature vector is determined. The determination layer outputs the missing determination result of the person under guardianship based on the second feature vector. By making a missing determination based on the second feature vector generated from the multiple historical feature vectors and the first feature vector, the efficiency and accuracy of missing determination can be effectively improved.

[0089] It should be noted that the historical feature vector is determined based on the second data to be detected, and the acquisition time of the second detection data is before the first detection data. For example, if the first detection data is collected at 10 o'clock, the second detection data can be collected at 9:55.

[0090] In some embodiments, the method for determining the second feature vector according to the first feature vector and multiple historical feature vectors may be: determining the similarity between the first feature vector and each historical feature vector; constructing a similarity matrix according to the similarity between the first feature vector and each historical feature vector; encoding the similarity matrix through an encoder to obtain the second feature vector. By constructing a similarity matrix for the first feature vector and multiple historical feature vectors and encoding the similarity matrix, the second feature vector can be accurately obtained.

[0091] In some embodiments, the method for constructing a similarity matrix according to the similarity between the first feature vector and each historical feature vector may be: the method for determining the similarity between the first feature vector and each historical feature vector may be: obtaining a preset similarity formula, and the preset similarity formula is The s i,j is the similarity, the h i,k is the first feature vector, and the h j,k is the historical feature vector. Based on the preset similarity formula, the first feature vector and the historical feature vector are calculated to obtain a similarity; and based on the preset similarity formula, the first feature vector and each historical feature vector are calculated respectively to obtain multiple similarities.

[0092] In some embodiments, the method for constructing a similarity matrix according to the similarity between the first feature vector and each historical feature vector may be: arranging the similarities according to the time sequence of the time points generated by the historical feature vectors corresponding to the similarities to obtain a similarity matrix. By arranging the similarities according to the time sequence of the time points generated by the historical feature vectors, the similarity matrix can be accurately obtained.

[0093] In some embodiments, the method for outputting the predicted wandering determination result based on the first feature vector of the determination layer may be as follows: performing feature calculation on the first feature vector through a preset prediction classification output formula to obtain a classification output feature value; calculating the classification output feature value through a preset wandering evaluation formula to obtain a wandering evaluation score; if the wandering evaluation score is greater than or equal to a preset score, determining that the wandering determination result is that the person under guardianship has wandered. Herein, the preset score may be set according to actual circumstances, and the embodiments of the present application do not make specific limitations thereto. For example, the preset score may be set to 60%. By processing the first feature vector through the determination layer, the predicted wandering determination result can be accurately obtained.

[0094] Exemplarily, the preset wandering evaluation formula is where E is the classification output feature value, and f k is the classification output function. Herein, the present application takes the K-classification as an example for illustration. Of course, the f k may also be other types of functions. Based on the preset wandering evaluation formula, the classification output feature value is calculated to obtain the classification output feature value.

[0095] Exemplarily, the preset wandering evaluation formula is where y i is the predicted wandering evaluation score, a i is the classification output feature value, and a j is the classification output feature value. It can be understood that the denominator of the preset wandering evaluation formula is the sum of multiple classification output feature values. Based on the preset wandering evaluation formula, the classification output feature value is calculated to obtain the wandering evaluation score.

[0096] In some embodiments, performing feature calculation on the second feature vector through a preset prediction classification output formula to obtain a classification output feature value; calculating the classification output feature value through a preset wandering evaluation formula to obtain a wandering evaluation score; if the wandering evaluation score is greater than or equal to a preset score, determining that the wandering determination result is that the person under guardianship has wandered.

[0097] Exemplarily, in the scenario where an elderly person checks into a nursing home, the elderly person wears a smart wearable device. The smart wearable device obtains first data to be detected in the environment where the elderly person is located. The first data to be detected includes first position data, first image data, and first voice data. The feature extraction network in a preset missing person discrimination model is used to respectively extract features from the first position data, the first image data, and the first voice data to obtain position features, image features, and voice features. The feature fusion network in the missing person discrimination model is used to perform fusion processing on the position features, the image features, and the voice features to obtain fused features. The encoder in the missing person discrimination model is used to perform encoding processing on the fused features to obtain a first feature vector. The determination layer in the missing person discrimination model outputs a missing person determination result of the elderly person based on the first feature vector. By processing the first detection data of the elderly person, it is possible to accurately determine whether the elderly person is missing, greatly improving the safety of the elderly person and the service quality of the nursing home.

[0098] Exemplarily, in the field of medical and health, in medical data statistics, it is necessary to count the number of times whether there is a phenomenon of a patient getting lost in daily life. The patient wears a smart wearable device, and the first data to be detected in the environment where the patient is located is obtained through the smart wearable device at preset time intervals. The first data to be detected includes first position data, first image data, and first voice data. The feature extraction network in a preset missing person discrimination model is used to respectively extract features from the first position data, the first image data, and the first voice data to obtain position features, image features, and voice features. The feature fusion network in the missing person discrimination model is used to perform fusion processing on the position features, the image features, and the voice features to obtain fused features. The encoder in the missing person discrimination model is used to perform encoding processing on the fused features to obtain a first feature vector. The determination layer in the missing person discrimination model outputs a missing person determination result of the patient based on the first feature vector. By processing the first detection data of the patient, it is possible to accurately determine whether the elderly person is missing, thereby greatly improving the accuracy of data statistics and providing reliable data for subsequent disease treatment.

[0099] In some embodiments, if the missing person determination result is that the person under guardianship is missing, the first position data and the first image data are sent to the guardian corresponding to the person under guardianship, so that the guardian can find the person under guardianship in a timely manner according to the first position data and the first image data. In the case where the person under guardianship is missing, the first position data and the first image data are transmitted to the guardian, enabling the guardian to retrieve the person under guardianship in a timely manner, thereby reducing the risk of accidents occurring to the person under guardianship.

[0100] The method for determining the loss of a person provided in the above embodiment obtains first data to be detected in the environment where the person under guardianship is located, and the first data to be detected includes first position data, first image data, and first voice data; respectively extracts features from the first position data, the first image data, and the first voice data through a feature extraction network in a preset person loss determination model to obtain position features, image features, and voice features. The person loss determination model is obtained by pre-training a multi-modal neural network model based on multiple training samples, and the training samples include sample position data, sample image data, sample voice data, and labeled loss determination results; performs fusion processing on the position features, the image features, and the voice features through a feature fusion network in the person loss determination model to obtain fusion features; performs encoding processing on the fusion features through an encoder in the person loss determination model to obtain a first feature vector; and outputs a loss determination result of the person under guardianship based on the first feature vector through a determination layer in the person loss determination model. In this application, by identifying the first position data, the first image data, and the first voice data in the environment where the person under guardianship is located, the environmental perception ability can be improved, thereby improving the efficiency and accuracy of loss judgment. Then, based on the multi-modal neural network model, loss judgment is performed on the first position data, the first image data, and the first voice data. The multi-modal neural network model can improve the accuracy of data processing, thereby greatly improving the efficiency and accuracy of loss determination.

[0101] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a device for determining the loss of a person provided in an embodiment of the present application.

[0102] As Figure 4 shown, the person loss determination device 300 includes an acquisition module 310, a feature extraction module 320, a fusion processing module 330, an encoding processing module 340, and an output module 350, where:

[0103] The acquisition module 310 is configured to acquire first data to be detected in the environment where the person under guardianship is located, and the first data to be detected includes first position data, first image data, and first voice data;

[0104] The feature extraction module 320 is configured to respectively extract features from the first position data, the first image data, and the first voice data through a feature extraction network in a preset person loss determination model to obtain position features, image features, and voice features. The person loss determination model is obtained by pre-training a multi-modal neural network model based on multiple training samples, and the training samples include sample position data, sample image data, sample voice data, and labeled loss determination results;

[0105] The fusion processing module 330 is configured to perform fusion processing on the location feature, the image feature, and the voice feature through the feature fusion network in the missing person discrimination model to obtain a fusion feature;

[0106] The encoding processing module 340 is configured to perform encoding processing on the fusion feature through the encoder in the missing person discrimination model to obtain a first feature vector;

[0107] The output module 350 is configured to output a missing determination result of the person under guardianship based on the first feature vector through a determination layer in the missing person discrimination model.

[0108] In some embodiments, the output module 350 is further configured to:

[0109] Obtain a plurality of historical feature vectors, where the historical feature vectors are determined according to second data to be detected, and the acquisition time of the second data to be detected is before the first data to be detected;

[0110] Determine a second feature vector according to the first feature vector and the plurality of historical feature vectors;

[0111] The outputting the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the missing person discrimination model includes:

[0112] Outputting the missing determination result of the person under guardianship based on the second feature vector through the determination layer.

[0113] In some embodiments, the output module 350 is further configured to:

[0114] Determine the similarity between the first feature vector and each of the historical feature vectors;

[0115] Construct a similarity matrix according to the similarity between the first feature vector and each of the historical feature vectors;

[0116] Encode the similarity matrix through the encoder to obtain the second feature vector.

[0117] In some embodiments, the output module 350 is further configured to:

[0118] Arrange the similarities according to the time sequence of the time points corresponding to the historical feature vectors generated by each of the similarities to obtain a similarity matrix.

[0119] In some embodiments, the output module 350 is further configured to:

[0120] Perform feature calculation on the first feature vector through a preset prediction classification output formula to obtain a classification output feature value;

[0121] Calculate the classification output feature value through a preset missing evaluation formula to obtain a missing evaluation score;

[0122] If the missing evaluation score is greater than or equal to a preset score, determine that the missing determination result is that the person under guardianship is missing.

[0123] Please refer to Figure 5 , FIG. 5 is a schematic block diagram of another personnel missing discrimination device provided by an embodiment of the present application.

[0124] As Figure 5 shown, the personnel missing discrimination device 400 includes an acquisition module 410, a selection module 420, a training module 430, a determination module 440, and a generation module 450, where:

[0125] The acquisition module 410 is configured to acquire a sample data set, the sample data set includes a plurality of sample data, and the sample data includes sample position data, sample image data, sample voice data, and a labeled missing determination result;

[0126] The acquisition module 410 is further configured to acquire a preset multi-modal neural network model;

[0127] The selection module 420 is configured to select a sample data from the sample data set as the target sample data;

[0128] The training module 430 is configured to input the sample position data, sample image data, and sample voice data in the target sample data into the preset multi-modal neural network model for training to obtain a predicted missing determination result;

[0129] The determination module 440 is configured to determine whether the preset multi-modal neural network model converges according to the predicted missing determination result and the labeled missing determination result;

[0130] The generation module 450 is configured to, if the preset multi-modal neural network model does not converge, continue to execute the step of selecting a sample data from the sample data set as the target sample data until the preset multi-modal neural network model converges to obtain a personnel missing discrimination model.

[0131] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above personnel missing discrimination device can refer to the corresponding process in the foregoing embodiment of the personnel missing discrimination method, and will not be elaborated herein.

[0132] Please refer toFigure 6 , Figure 6 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.

[0133] As Figure 6 shown, the computer device 500 includes a processor 502 and a memory 503 connected through a system bus 501. Among them, the memory 503 may include a storage medium and an internal memory.

[0134] The storage medium can store a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any method for judging the loss of a person.

[0135] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device.

[0136] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any method for judging the loss of a person.

[0137] Those skilled in the art can understand that Figure 6 the structure shown in

[0138] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0139] Among them, in one embodiment, the processor 502 is used to run the computer program stored in the memory to implement the following steps:

[0140] Obtain first detection data of the environment where the person under guardianship is located, and the first detection data includes first position data, first image data, and first voice data;

[0141] The feature extraction network in the preset person missing discrimination model is used to extract features from the first location data, the first image data, and the first voice data respectively, to obtain location features, image features, and voice features. The person missing discrimination model is obtained by pre-training a multi-modal neural network model based on multiple training samples, and the training samples include sample location data, sample image data, sample voice data, and labeled missing determination results;

[0142] The feature fusion network in the person missing discrimination model is used to perform fusion processing on the location features, the image features, and the voice features to obtain fusion features;

[0143] The encoder in the person missing discrimination model is used to perform encoding processing on the fusion features to obtain a first feature vector;

[0144] The determination layer in the person missing discrimination model is used to output the missing determination result of the person under guardianship based on the first feature vector.

[0145] In one embodiment, before the processor 502 realizes outputting the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model, it is further used to realize:

[0146] Obtain multiple historical feature vectors, where the historical feature vectors are determined according to the second data to be detected, and the acquisition time of the second data to be detected is before the first data to be detected;

[0147] Determine a second feature vector according to the first feature vector and the multiple historical feature vectors;

[0148] The outputting the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model includes:

[0149] Outputting the missing determination result of the person under guardianship based on the second feature vector through the determination layer.

[0150] In one embodiment, when the processor 502 realizes determining the second feature vector according to the first feature vector and the multiple historical feature vectors, it is used to realize:

[0151] Determine the similarity between the first feature vector and each historical feature vector;

[0152] Construct a similarity matrix according to the similarity between the first feature vector and each historical feature vector;

[0153] Encode the similarity matrix through the encoder to obtain the second feature vector.

[0154] In one embodiment, when the processor 502 constructs a similarity matrix according to the similarity between the first feature vector and each of the historical feature vectors, it is configured to:

[0155] Arrange the similarities based on the chronological order of the time points corresponding to the historical feature vectors corresponding to the similarities, to obtain a similarity matrix.

[0156] In one embodiment, when the processor 502 outputs a lost determination result of the person under guardianship based on the first feature vector through a determination layer in the person lost determination model, it is configured to:

[0157] Perform feature calculation on the first feature vector through a preset prediction classification output formula to obtain a classification output feature value;

[0158] Calculate the classification output feature value through a preset lost evaluation formula to obtain a lost evaluation score;

[0159] If the lost evaluation score is greater than or equal to a preset score, determine that the lost determination result is that the person under guardianship is lost.

[0160] In one embodiment, the processor 502 is further configured to:

[0161] Obtain a sample data set, where the sample data set includes a plurality of sample data, and the sample data includes sample location data, sample image data, sample voice data, and an annotated lost determination result;

[0162] Obtain a preset multi-modal neural network model, and select a sample data from the sample data set as target sample data;

[0163] Input the sample location data, sample image data, and sample voice data in the target sample data into the preset multi-modal neural network model for training to obtain a predicted lost determination result;

[0164] Determine whether the preset multi-modal neural network model converges according to the predicted lost determination result and the annotated lost determination result;

[0165] If the preset multi-modal neural network model does not converge, continue to execute the step of selecting a sample data from the sample data set as target sample data until the preset multi-modal neural network model converges to obtain a person lost determination model.

[0166] In one embodiment, after the processor 502 inputs the missing person prediction feature value into the determination layer to determine whether a person is missing and obtains a missing person determination result, the processor 502 is further configured to:

[0167] If the missing person determination result indicates that the person under guardianship is missing, send the first location data and the first image data to the guardian corresponding to the person under guardianship, so that the guardian can search for the person under guardianship in a timely manner according to the first location data and the first image data.

[0168] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described computer device can refer to the corresponding process in the foregoing embodiment of the method for determining missing persons, and will not be elaborated herein.

[0169] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed can refer to the various embodiments of the method for determining missing persons in the present application.

[0170] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may be non-volatile or volatile. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device.

[0171] Further, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node.

[0172] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. A blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0173] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0174] It should also be understood that the term "and / or" used in the specification of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or system comprising the element.

[0175] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A method for discriminating the loss of a person based on artificial intelligence, characterized in that, Including: Obtain first data to be detected of the environment where the person under guardianship is located, where the first data to be detected includes first position data, first image data, and first voice data; Respectively perform feature extraction on the first position data, the first image data, and the first voice data through a feature extraction network in a preset person missing discrimination model. Obtain a position feature, an image feature, and a voice feature. The person missing discrimination model is pre-trained on a multi-modal neural network model based on multiple training samples. The training samples include sample position data, sample image data, sample voice data, and labeled missing determination results; Perform fusion processing on the position feature, the image feature, and the voice feature through a feature fusion network in the person missing discrimination model to obtain a fusion feature; Perform encoding processing on the fusion feature through an encoder in the person missing discrimination model to obtain a first feature vector; Output a missing determination result of the person under guardianship based on the first feature vector through a determination layer in the person missing discrimination model.

2. The method for determining the loss of a person according to claim 1, wherein Before outputting the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model, it further includes: Obtain multiple historical feature vectors, where the historical feature vectors are determined according to second data to be detected, and the collection time of the second data to be detected is before the first data to be detected; Determine a second feature vector according to the first feature vector and the multiple historical feature vectors; Outputting the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model includes: Output the missing determination result of the person under guardianship based on the second feature vector through the determination layer.

3. The method for determining the loss of a person according to claim 2, characterized in that, Determining the second feature vector according to the first feature vector and the multiple historical feature vectors includes: Determine the similarity between the first feature vector and each historical feature vector; Construct a similarity matrix according to the similarity between the first feature vector and each historical feature vector; Encode the similarity matrix through the encoder to obtain the second feature vector.

4. The method for judging the loss of a person according to claim 3, wherein Constructing the similarity matrix according to the similarity between the first feature vector and each historical feature vector includes: Arrange the similarities according to the time sequence of the time points when the historical feature vectors corresponding to the similarities are generated to obtain a similarity matrix.

5. The method for determining the loss of a person according to claim 1, characterized in that, Outputting the missing determination result of the person under guardianship based on the first feature vector through the determination layer in the person missing discrimination model includes: Perform feature calculation on the first feature vector through a preset prediction classification output formula to obtain a classification output feature value; Calculate the missing evaluation score through a preset missing evaluation formula for the classification output feature value; If the missing evaluation score is greater than or equal to a preset score, determine that the missing determination result is that the person under guardianship is missing.

6. The method for judging the loss of a person according to any one of claims 1-5, characterized in that The method further includes: Obtain a sample data set, where the sample data set includes multiple sample data, and the sample data includes sample location data, sample image data, sample voice data, and an annotated result of the missing person determination; Obtain a preset multi-modal neural network model, and select a sample data from the sample data set as the target sample data; Input the sample location data, sample image data, and sample voice data in the target sample data into the preset multi-modal neural network model for training to obtain a predicted result of the missing person determination; Determine whether the preset multi-modal neural network model converges according to the predicted result of the missing person determination and the annotated result of the missing person determination; If the preset multi-modal neural network model does not converge, continue to execute the step of selecting a sample data from the sample data set as the target sample data until the preset multi-modal neural network model converges to obtain a missing person discrimination model.

7. The method for judging the loss of a person according to any one of claims 1-5, characterized in that, After inputting the missing person prediction feature value into the determination layer to perform the missing person determination and obtaining the missing person determination result, it further includes: If the missing person determination result is that the ward is missing, send the first location data and the first image data to the guardian corresponding to the ward, so that the guardian can search for the ward in a timely manner according to the first location data and the first image data.

8. A device for judging the loss of a person, characterized in that, The missing person discrimination device includes an acquisition module, a feature extraction module, a fusion processing module, an encoding processing module, and an output module, where: The acquisition module is configured to acquire first data to be detected in the environment where the ward is located, and the first data to be detected includes first location data, first image data, and first voice data; The feature extraction module is configured to respectively perform feature extraction on the first location data, the first image data, and the first voice data through a feature extraction network in a preset missing person discrimination model to obtain a location feature, an image feature, and a voice feature. The missing person discrimination model is pre-trained based on multiple training samples for a multi-modal neural network model. The training samples include sample location data, sample image data, sample voice data, and an annotated result of the missing person determination; The fusion processing module is configured to perform fusion processing on the location feature, the image feature, and the voice feature through a feature fusion network in the missing person discrimination model to obtain a fusion feature; The encoding processing module is configured to perform encoding processing on the fusion feature through an encoder in the missing person discrimination model to obtain a first feature vector; The output module is configured to output the missing person determination result of the ward based on the first feature vector through a determination layer in the missing person discrimination model.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the missing person discrimination method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for judging the loss of a person according to any one of claims 1 to 7 are implemented.