Intelligent nursing system and method
Through a monocular camera and deep learning algorithm, a 3D reconstruction virtual picture that integrates the human body and scene is generated, which solves the privacy and security of the intelligent care system and data transmission and storage problems, and achieves efficient monitoring and nursing capabilities.
Patent Information
- Application Number
- CN202311223433.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-20
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-09-20
AI Technical Summary
The existing intelligent care system has problems with monitoring privacy security, data transmission and storage, and cannot cope with the care needs in complex environments. It lacks deep learning and complex artificial intelligence technologies, resulting in limited recognition and adaptability.
A monocular camera is used to combine deep learning algorithms and 3D reconstruction technology to extract and match semantics through a data analysis server, generate a 3D reconstruction virtual picture that integrates the human body and the scene, and use a pre-constructed character interaction knowledge graph to analyze the interaction status between the human body and the scene, and transmit a small amount of semantic information for monitoring edge reconstruction.
It improves the privacy and security of the monitoring system, reduces the data transmission bandwidth requirements and storage costs, enhances the system's identification accuracy and care quality in complex environments, and provides timely decision-making support.
Smart Images

Figure CN117315532B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent nursing technology, and in particular to an intelligent nursing system and method. Background Art
[0002] Traditional smart care systems refer to early or basic intelligent technologies used in the fields of healthcare and long-term care, with the main goal of providing auxiliary monitoring and care services. Although smart care systems have made certain progress in the past few years, there are still some significant shortcomings. Traditional smart care systems usually include basic sensors and data acquisition devices. These sensors and data acquisition devices are used to monitor patients' physiological parameters, activity levels, and environmental conditions. These systems can provide real-time monitoring and automatically alarm or notify medical staff so that they can respond in time when problems arise.
[0003] With the widespread application of artificial intelligence technology in real life, research on nursing systems based on machine learning algorithms is developing rapidly. Machine learning algorithms are used to analyze data and predict possible risk events. However, traditional intelligent nursing systems still have some limitations, which may limit their effectiveness and feasibility in practical applications. The specific limitations are mainly reflected in the following aspects:
[0004] (1) Privacy and security issues: Traditional smart care systems need to collect and process a large amount of patient data, such as patient video data, which may cause concerns about patient privacy and data security. If the system is attacked by hackers or data is leaked, it may pose serious risks to patients.
[0005] (2) Limited intelligence: Traditional intelligent care systems are often based on rules and simple machine learning algorithms, lacking deep learning and more complex artificial intelligence technologies. This limits the system's ability to identify and adapt to complex situations, making it difficult to truly achieve intelligent care.
[0006] (3) User data transmission and storage issues: Traditional care systems usually use video streaming to monitor and transmit data in real time when monitoring users. This has high requirements for network communication and greatly increases storage costs. Long-term operation will generate a lot of data maintenance costs.
[0007] Therefore, there is an urgent need to provide an intelligent care system that can improve the recognition and adaptability of the care system through deep learning and complex artificial intelligence technologies, and solve the problems of monitoring system privacy security, data transmission and storage in the existing care system. Summary of the Invention
[0008] The present invention provides an intelligent nursing system and method, which solves the technical problem that existing nursing systems have problems such as monitoring system privacy security, data transmission and storage, and are unable to cope with nursing problems in complex environments.
[0009] To solve the above technical problems, the present invention provides an intelligent nursing system and method.
[0010] In a first aspect, the present invention provides an intelligent nursing system applied to an intelligent monitoring system, wherein the intelligent monitoring system includes a data analysis server and multiple monitoring edge terminals and multiple visual perception terminals that establish communication connections with the data analysis server, wherein the visual perception terminal and the data analysis server are located in the same local area network, and the visual perception terminal includes a monocular camera. The system includes:
[0011] The visual perception terminal is used to capture the original nursing video stream through the monocular camera and upload it to the data analysis server;
[0012] The data analysis server is configured to perform semantic extraction on the original image frames in the original nursing video stream using a deep learning algorithm to obtain semantic information of different categories, and transmit the semantic information to the monitoring edge end via a preset network transmission protocol; wherein the semantic information includes human semantic information, scene semantic information, and human-scene interaction status information between the human body and the scene;
[0013] The monitoring edge is used to match different categories of semantic information with its stored three-dimensional model of people and scenes to obtain a semantic matching model, and load and drive the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual picture that integrates the human body and the scene.
[0014] In a further embodiment, the data analysis server is specifically used to:
[0015] A deep learning algorithm is used to extract human semantic information and scene semantic information from the original image frames of the original nursing video stream, and a pre-built human interaction knowledge graph is used to analyze the human semantic information and the scene semantic information to obtain the human-scene interaction status information between the human body and the scene.
[0016] In a further embodiment, the human body semantic information includes human body behavior state information and human body position information, and the human body semantic information is extracted from the original image frames of the original nursing video stream using a deep learning algorithm, specifically:
[0017] Using a pre-trained target detection model to identify the original image frame to obtain a human target detection frame;
[0018] Inputting the human target detection frame and the original image frame into an improved posture estimation model to obtain human skeleton point coordinate information; the improved posture estimation model is a posture estimation model embedded with a posture structured expression algorithm;
[0019] Inputting the human skeleton point coordinate information and the original image frame into a spatiotemporal graph convolutional network model to obtain character behavior state information;
[0020] The preset target skeleton point coordinate information is obtained from the human body skeleton point coordinate information, and the target skeleton point coordinate information is converted into the human body three-dimensional coordinates in the world coordinate system by using camera calibration to obtain the human body position information.
[0021] In a further embodiment, the semantic information further includes event information, and the event information includes normal event information and abnormal event information. The data analysis server is further configured to:
[0022] The event information is transmitted to the monitoring edge end through a preset network transmission protocol, so that the monitoring edge end monitors the event information and generates alarm information according to the abnormal event information when abnormal event information is detected.
[0023] In a further embodiment, the 3D reconstructed virtual screen is an interactive 3D reconstructed virtual screen, and the monitoring edge end is further used to:
[0024] When a model interaction trigger event is detected, in response to the model interaction trigger event, model identification data corresponding to the model interaction trigger event is obtained, and human-scene three-dimensional model information corresponding to the model identification data is retrieved and displayed according to the model identification data.
[0025] In a further embodiment, the human-scene 3D model includes a human body 3D model and a scene 3D model;
[0026] The human-scene three-dimensional model is stored in a three-dimensional model library at the monitoring edge end, and the three-dimensional model library is stored in a relational database management system in the form of a server path. The relational database management system includes an event management unit for storing the event information and a model management unit for storing all human-scene three-dimensional model information.
[0027] In a further embodiment, the preset network transmission protocol includes the User Datagram Protocol.
[0028] In a second aspect, the present invention provides an intelligent nursing method applied to an intelligent monitoring system, wherein the intelligent monitoring system includes a data analysis server and multiple monitoring edge terminals and multiple visual perception terminals that establish communication connections with the data analysis server, wherein the visual perception terminal and the data analysis server are located in the same local area network, and the visual perception terminal includes a monocular camera. The method includes the following steps:
[0029] Capturing and uploading the original nursing video stream through the monocular camera;
[0030] Using a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain different categories of semantic information; wherein the semantic information includes human semantic information, scene semantic information, and human-scene interaction state information between the human body and the scene;
[0031] Semantic information of different categories is matched with the stored three-dimensional model of human scene to obtain a semantic matching model, and the semantic matching model corresponding to each category of semantic information is loaded and driven to generate a 3D reconstructed virtual picture of the fusion of human body and scene.
[0032] In a further embodiment, the step of using a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain semantic information of different categories includes:
[0033] A deep learning algorithm is used to extract human semantic information and scene semantic information from the original image frames of the original nursing video stream, and a pre-built human interaction knowledge graph is used to analyze the human semantic information and the scene semantic information to obtain the human-scene interaction status information between the human body and the scene.
[0034] In a third aspect, the present invention further provides an electronic device having an intelligent nursing system as described above.
[0035] The present invention provides an intelligent nursing system and method. The system captures and uploads a raw nursing video stream via a monocular camera. A data analysis server uses a deep learning algorithm to perform semantic extraction on the raw image frames in the raw nursing video stream to obtain different categories of semantic information. This semantic information is then transmitted to a monitoring edge terminal via a preset network transmission protocol, so that the monitoring edge terminal matches the different categories of semantic information with its stored three-dimensional model of a person and a scene to obtain a semantic matching model. The monitoring edge terminal then loads and drives the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual image that fuses the person and the scene. Compared to existing technologies, the present invention reconstructs information about the monitored person and the scene at the monitoring edge terminal by transmitting a small amount of semantic information. The reconstructed information only provides the status and location of the monitored person and the scene, thereby improving the privacy and security of the monitoring system. Furthermore, by processing the nursing video using a deep learning algorithm, the system can handle nursing videos in various complex environments, improving recognition accuracy and nursing quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a block diagram of an intelligent nursing system provided by an embodiment of the present invention;
[0037] Figure 2 Schematic diagram of the structure of a traditional monitoring system provided by an embodiment of the present invention;
[0038] Figure 3 This is a schematic diagram of the structure of the intelligent monitoring system provided by an embodiment of the present invention;
[0039] Figure 4 This is an example diagram of the character interaction knowledge graph provided by an embodiment of the present invention;
[0040] Figure 5 This is an example diagram of an application of the intelligent nursing method provided by an embodiment of the present invention;
[0041] Figure 6 This is a traditional 3D reconstruction example diagram provided by an embodiment of the present invention;
[0042] Figure 7 This is an example diagram of a 3D reconstructed virtual image provided by an embodiment of the present invention;
[0043] Figure 8 It is a flowchart of the intelligent care method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.
[0045] refer to Figure 1 An embodiment of the present invention provides an intelligent nursing system, which is applied to an intelligent monitoring system. The intelligent monitoring system includes a visual perception terminal 101, a data analysis server 102 and a monitoring edge terminal 103. The data analysis server 102 establishes a communication connection with the multiple monitoring edge terminals 103 and the multiple visual perception terminals 102, wherein the visual perception terminal and the data analysis server are located in the same local area network, and the visual perception terminal 103 includes a monocular camera.
[0046] In this embodiment, the visual perception terminal 101 is used to capture and upload the original nursing video stream to the data analysis server through the monocular camera, wherein the original nursing video stream includes multiple original image frames.
[0047] It should be noted that, in this embodiment, the user can deploy the visual perception terminal and the data analysis server at a specific location, and make the two in the same local area network and start the camera and the server at the same time. In this embodiment, the RTSP stream address of the monocular camera is input into the data analysis server, wherein RTSP (Real-Time Streaming Protocol) is usually used for streaming media applications. It is a network protocol for controlling real-time media streams. RTSP is usually used to transmit audio, video and other streaming media data on a computer network. In this embodiment, the real-time streaming protocol RTSP is used to transmit the video stream captured by the network camera of the visual perception terminal to the data analysis server. After receiving the transmitted video stream, the data analysis server uses the OpenCV computer vision library to decompose the video into original image frames in a specific RTSP stream. In this embodiment, a monocular camera is selected as the network camera for video stream acquisition, such as Figure 2 、 Figure 3 As shown, compared with the traditional monitoring system structure, this embodiment does not need to transmit data separately for remote monitoring and intranet monitoring, which saves a lot of bandwidth compared to traditional monitoring communication transmission. Taking the network camera used in the embodiment of the present invention as an example, the size of a frame of 1280*720 resolution image is about 1MB, and the corresponding semantic information is compressed to less than 500 bytes.
[0048] Traditional nursing systems use special cameras such as depth cameras and infrared cameras with depth information, or increase the number of cameras and install multiple cameras at various angles to collect videos or images. The embodiments of the present invention are based on a monocular camera combined with camera calibration and other technologies to judge the scene status information and position of the characters. It should be noted that although traditional human body recognition also uses a monocular camera to sample human posture information, due to the small amount of information collected by the monocular camera, the sampled information cannot be used to make a digital twin. This application combines deep learning algorithms and 3D scene reconstruction methods to achieve the effect of digital twins.
[0049] In this embodiment, the data analysis server 102 includes a semantic extraction module, which is used to use a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain semantic information of different categories, and transmit the semantic information to the monitoring edge through a preset network transmission protocol; wherein, the semantic information includes human body semantic information, scene semantic information and human-scene interaction status information between the human body and the scene; the scene semantic information includes scene object position information; the preset network transmission protocol includes the User Datagram Protocol.
[0050] The semantic extraction module is specifically used to extract human semantic information and scene semantic information from the original image frames of the original nursing video stream using a deep learning algorithm, and analyze the human semantic information and scene semantic information using a pre-built human interaction knowledge graph to obtain the human-scene interaction state information between the human body and the scene, wherein the human semantic information includes human behavior state information and human position information. The extraction of human semantic information from the original image frames of the original nursing video stream using a deep learning algorithm is specifically as follows:
[0051] Using a pre-trained target detection model to identify the original image frame to obtain a human target detection frame;
[0052] Inputting the human target detection frame and the original image frame into an improved posture estimation model to obtain coordinate information of human skeleton points; the improved posture estimation model is a posture estimation model embedded with a posture structured expression algorithm; the posture structured expression algorithm is used to predict human skeleton points;
[0053] Inputting the human skeleton point coordinate information and the original image frame into a spatiotemporal graph convolutional network model to obtain character behavior state information;
[0054] The preset target skeleton point coordinate information is obtained from the human body skeleton point coordinate information, and the target skeleton point coordinate information is converted into the human body three-dimensional coordinates in the world coordinate system by using camera calibration to obtain the human body position information.
[0055] Specifically, the data analysis server 102 starts three threads during operation, wherein the first thread is used to capture the original monitoring video stream pushed by the monocular camera through the RTSP protocol, and send the original monitoring video stream in the form of frames into a shared queue; the second thread is used to continuously extract the latest stored original image frames from the shared queue to ensure the lowest delay, and continuously send the latest stored original image frames into the semantic extraction module; the third thread is used to continuously transmit the results extracted by the semantic extraction module to the corresponding monitoring edge end through the UDP protocol (User Datagram Protocol).
[0056] After receiving the raw image frames continuously passed in by the second thread, the semantic extraction module uses a trained neural network combined with camera imaging principles to perform semantic extraction. The extracted semantic information includes human behavior status information, human position information, scene semantic information, and human-scene interaction status information. Among them, scene semantic information includes object position information in the scene, and object position information in the scene includes the position information of the bed, chair, table, sofa, and water cup. The semantic extraction module performs semantic extraction as follows:
[0057] Regarding the process of extracting human behavior state information, this embodiment inputs the original image frame into a pre-trained target detection model for recognition. The target detection model extracts a target detection frame of the human category from the original image frame. The target detection model includes but is not limited to a general target detection model such as the YOLO algorithm, the SSD algorithm, and the RCNN algorithm. Then, the target detection frame output by the target detection model and the original image frame are input into a pre-trained improved pose estimation model to obtain the pixel coordinates of the human skeleton point. It should be noted that compared with network camera devices such as depth cameras, this embodiment uses a monocular camera to capture video to reduce costs. However, the difficulty of capturing semantic information with a monocular camera compared to a depth camera lies in the lack of depth information and the scale ambiguity of scene objects. At the same time, compared with multi-cameras, monocular cameras lack other perspective information. Under occlusion conditions, a large amount of feature information of the monitored object is also lost. Therefore, in order to solve the problem of occlusion in the captured image caused by the limitations of the monocular camera, this embodiment adopts the PCT (Human Pose as Compositional Tokens) algorithm and embeds the PCT (Human Pose as Compositional Tokens) in the pose estimation model. An improved posture estimation model is constructed by using a posture tokens (Structured Expression of Posture) algorithm to predict skeleton points using the PCT algorithm to obtain complete pixel coordinate information of human skeleton points. The PCT algorithm adopts a top-down approach to complete the detected human skeleton points, ensuring that the improved posture estimation model outputs skeleton point information completely and accurately even in the case of occlusion. Then, this embodiment inputs the human skeleton point pixel coordinate information and the original image frame output by the improved posture estimation model into the trained spatiotemporal graph convolutional network model STGCN to obtain character behavior status information, including but not limited to standing, sitting, lying, walking and other information.
[0058] For the target position extraction process detected, this embodiment uses camera calibration to convert the pixel coordinates in the image into world coordinates. Through the internal parameters and external parameters of the camera, the internal parameters include focal length, principal point position, etc., and the external parameters include the position and direction of the camera in the world coordinate system. This embodiment obtains the preset target skeleton point coordinate information from the moral human skeleton point coordinate information of the previous step. The target skeleton point coordinate information includes the pixel coordinates of the ankle in the human skeleton point. The pixel coordinates of the ankle in the human skeleton point are converted to obtain the real-world three-dimensional coordinates of the human body. The specific principle can refer to the Zhang Zhengyou calibration method camera calibration.
[0059] This embodiment transmits the original video stream captured by the monocular camera directly to the semantic extraction module. The extracted semantic information does not contain the privacy of the user being monitored, but only expresses "what the user did in what time and scene", so that when the monitoring edge uses the 3D engine to reconstruct the transmitted semantic information, the privacy details are visually filtered out. At the same time, the database only stores semantic information, and the video stream captured by the camera will be released from the computer memory after the semantic information is extracted, and no record will be made. At the same time, in addition to pure video care, traditional nursing systems also include a nursing method that relies on wearing monitoring electronic devices to judge the status of the monitored person. This type of monitoring system protects the user's privacy but has very little information. The guardian can only receive the ward's own information and ignore the environmental information, which may lead to inaccurate or missing abnormal judgments. For example, it is easy to confuse the two scenarios of lying in bed and fainting on the floor; on the other hand, the traditional wearable device monitoring method will also ignore the record of normal periodic behavior. For example, the ward is required to take medicine three times a day, and the wearable device monitoring can only remind the ward, and then the ward manually conveys the information of completing the task to the guardian. In order to solve this technical problem, this embodiment combines the knowledge graph to generate human-scene interaction status information, so as to extract the interaction between people and objects, and accurately and completely transmit the correct information without directly exposing the privacy of the ward. For example, when it is received that the human body is in a "lying" state, it is judged whether the human body is normal based on whether the human body interacts with the bed; on the other hand, the system database stores the interaction information between the human body and objects. Based on this information, the normal periodic behavior of the ward can be recorded for the guardian to view.
[0060] This embodiment combines the knowledge graph to realize the extraction process of semantic information of human-scene interaction status information, such as Figure 4As shown, this embodiment stores the predefined prior knowledge of human-object interaction in the form of a graph to form a human-object interaction knowledge graph. Therefore, when the semantic information of people and objects is transmitted, this embodiment combines the constructed human-object interaction knowledge graph to analyze and derive the interaction semantics between the two. From the perspective of semantic reconstruction, using the constructed human-object interaction knowledge graph to analyze and derive the interaction semantics between the two can improve the reconstruction accuracy and realism. For example, when a person sits on a chair in reality, the position and orientation information output by the algorithm may have errors, and the reconstructed image may show that the human body is sitting in the air. However, if the interaction between people and objects can be combined, it can be determined which person is interacting with which object, and then the semantics can be corrected at the reconstruction end. , and analyze that the person is sitting on the corresponding chair. At the same time, in terms of the accuracy of the recognition algorithm, this embodiment uses the constructed character interaction knowledge graph to analyze and derive the interaction semantics between the two, which increases the accuracy of distinguishing positive and abnormal behaviors. For example, a person slowly falls down when suffering from hypoglycemia, which is basically the same as the action of slowly lying down to go to bed. For such situations, even STGCN with timing information cannot make a judgment. Therefore, based on the STGCN model, this embodiment combines the information of the scene and can accurately distinguish the action of a person slowly falling down when suffering from hypoglycemia and slowly lying down to go to bed. For example, if a person has such an action on the bed or sofa, it can be judged that he is lying down to sleep. If the above action occurs in other places, it is judged as fainting.
[0061] In this embodiment, the monitoring edge end 103 is used to match different categories of semantic information with its stored human-scene 3D model to obtain a semantic matching model, and load and drive the semantic matching model corresponding to each category of semantic information at the monitoring edge end to generate a 3D reconstructed virtual image that integrates the human body and the scene. Figure 5 This is an example diagram of the application of the intelligent care method provided by an embodiment of the present invention.
[0062] Specifically, the monitoring edge end 103 starts two threads during operation. One thread continuously receives the semantic information of people and scenes, decodes and analyzes it, and puts the results into a shared queue; the other thread continuously takes out the latest result from the shared queue, reconstructs the semantic information in the result on the monitoring end, and stores the corresponding event information in the database. It should be noted that the database of the monitoring edge end is built based on MySQL, which is a relational database management system. The relational database management system includes an event management unit for storing the event information and a model management unit for storing all three-dimensional model information of the scene. The event management unit includes a form, which is connected through the connector class of Python in this embodiment. The access operations to the database are integrated through pre-written SQL statements, and the port is opened for the front-end to call. The calling process is that the front-end sends different operation requests to the database according to different ports, executes specific functions according to the requests sent, and executes specific SQL statements in the function.
[0063] The 3D human-scene model includes a 3D human model and a 3D scene model. The 3D human-scene model is stored in a 3D model library at the monitoring edge. The 3D model library is stored in a relational database management system in the form of a server path. In this embodiment, after decoding the semantic information, classified information is obtained. Semantic information of different categories is matched with models in the 3D model library. For example, if the semantic information sent includes a person, the "person" model in the model library is loaded. After the model corresponding to each category of semantic information received is loaded into the monitoring end display interface, the model is driven according to other semantic information under each category. For example, if the parsed semantic information expresses "person 1 sitting on chair 2", the monitoring edge will reconstruct the 3D model into the model corresponding to person 1 sitting on the model corresponding to chair 2. The semantic information also includes the person's movements and the interaction between the person and the object. This series of semantic reconstruction is completed by the animation library in the model. The reconstructed result is a 3D virtual image including the person and the scene. Compared with the monocular reconstruction scene, the advantage of this method is that even if there is occlusion in reality, the scene information of the person and the object can still be fully displayed.
[0064] In a specific embodiment, the monitoring edge end in this embodiment can be displayed on the web end, using the browser as the carrier, and using threejs to load and drive the model, wherein threejs is a WebGL third-party library written in Javascript, and is a 3D engine running in the browser; the front-end functional interface is developed using the Vue framework, which is a JavaScript framework for building user interfaces, and the 3D reconstructed virtual screen is an interactive 3D reconstructed virtual screen, that is, when the user interacts with the reconstructed 3D model, he can select or operate the reconstructed 3D model by clicking or touching to obtain the basic information of the model. Information, such as location, category, etc., the specific implementation process includes: when the monitoring edge end detects a model interaction trigger event, it responds to the model interaction trigger event, obtains the model identification data corresponding to the model interaction trigger event, and retrieves and displays the human-scene three-dimensional model information corresponding to the model identification data according to the model identification data. For example: when the user clicks on the model in the 3D reconstructed virtual picture, the listener brought by the Vue framework will listen to the mouse click event, and return the ID number of the model clicked by the mouse to the back end. The back end retrieves the information of the corresponding model from the model information table in the database according to the ID number of the model and sends it to the front end. The front end displays the form information on the display interface.
[0065] It should be noted that the result of traditional 3D reconstruction seems to include the scene, but in fact it is only a 3D reconstruction for the target of "people". That is to say, traditional 3D reconstruction only adds the information of people to the reconstruction, and the background is just a space composed of 2D textures. It is impossible to drag the viewing angle of the entire reconstructed scene. The embodiment of the present invention performs 3D reconstruction based on the semantic information of people, scenes, and the interaction between people and scenes. From the effect, it can be said to be a "digital twin". However, the difference from the existing digital twin is that the current digital twin generally detects the position and status of the person being monitored by wearing electronic devices, or expensive multi-eye cameras and depth cameras, while this embodiment uses a lower-cost monocular camera to achieve 3D scene reconstruction, etc. Figure 6 For traditional 3D reconstruction, it only adds human information to the reconstruction, and uses textures for the background. Figure 7 This is an example diagram of a 3D reconstructed virtual image provided by an embodiment of the present invention.
[0066] In this embodiment, the semantic information also includes various event information, and the event information includes normal event information and abnormal event information, for example: "Person 1 enters monitoring scene 1 at a certain time" is normal event information, and "Person 2 falls in monitoring scene 2" is abnormal event information. Both normal event information and abnormal event information will be stored in the corresponding database table according to time and place. The stored event information can be viewed and searched at the monitoring edge. After the data analysis server transmits the event information to the monitoring edge through a preset network transmission protocol, the monitoring edge monitors the event information, and when abnormal event information is detected, records the abnormal event information and generates an alarm information based on the abnormal event information to notify the monitoring personnel; when normal event information is detected, only the normal event information needs to be recorded, for example: the listener in the vue framework listens to the semantic information continuously transmitted. Once the transmitted semantic information contains abnormal event information, an alarm will be triggered, an alarm window will pop up, and the prepared alarm audio will be played. The alarm will be lifted after the monitoring personnel clicks the "OK" button.
[0067] Compared with the traditional nursing system that exposes the privacy of the monitored party to the monitoring party without reservation, the embodiment of the present invention reconstructs the information of the monitored party and the scene in the monitoring party through a small amount of transmitted semantic information. The reconstructed information only provides the status and location of the monitored party and the scene, so that the private information is filtered out, solving the privacy exposure problem of data transmission. At the same time, this embodiment compresses the transmitted video information, extracts its semantics and compresses it, and reconstructs the semantic information after receiving it at the monitoring end, thereby solving the high bandwidth problem required for traditional video transmission and the storage space overhead required for information log storage.
[0068] An embodiment of the present invention provides an intelligent care system that captures and uploads raw care video streams using a monocular camera to a data analysis server. The data analysis server uses a deep learning algorithm to extract semantic information from the raw image frames in the raw care video stream, obtaining different categories of semantic information. This information is then transmitted to a monitoring edge device via a pre-set network transmission protocol. The monitoring edge device then matches the different categories of semantic information with a stored three-dimensional model of a person and a scene to obtain a semantic matching model. The monitoring edge device then loads and drives the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual image that fuses the person and the scene. Compared to traditional care systems, the intelligent care system proposed in this embodiment not only enables accurate predictions, providing caregivers with more scientific and timely decision-making support, and helping to prevent and address various potential caregiving issues, but also reconstructs information about the monitored person and the scene within the monitoring device using a small amount of transmitted semantic information. This prevents the monitored person's privacy from being exposed to the monitoring device, saves storage space, and improves care efficiency and user experience.
[0069] In one embodiment, Figure 8 As shown, an embodiment of the present invention provides an intelligent care method, which is applied to an intelligent monitoring system. The intelligent monitoring system includes a data analysis server and multiple monitoring edge terminals and multiple visual perception terminals that establish communication connections with the data analysis server. The visual perception terminal and the data analysis server are located in the same local area network, and the visual perception terminal includes a monocular camera. The method includes the following steps:
[0070] S1. Capture and upload the original care video stream through the monocular camera;
[0071] S2. Using a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain different categories of semantic information; wherein the semantic information includes human semantic information, scene semantic information, and human-scene interaction status information between the human body and the scene;
[0072] S3. Match different categories of semantic information with the stored three-dimensional model of the human scene to obtain a semantic matching model, and load and drive the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual picture that integrates the human body and the scene.
[0073] In this embodiment, the step of using a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain semantic information of different categories includes:
[0074] A deep learning algorithm is used to extract human semantic information and scene semantic information from the original image frames of the original nursing video stream, and a pre-built human interaction knowledge graph is used to analyze the human semantic information and the scene semantic information to obtain the human-scene interaction status information between the human body and the scene.
[0075] It should be noted that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0076] For the specific definition of an intelligent nursing method, please refer to the above-mentioned definition of an intelligent nursing system, which will not be repeated here. Those skilled in the art will appreciate that the various modules and steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware, software, or a combination of both. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0077] An embodiment of the present invention provides an intelligent caregiving method. The method captures and uploads a raw caregiving video stream using a monocular camera; uses a deep learning algorithm to perform semantic extraction on raw image frames in the raw caregiving video stream to obtain different categories of semantic information; matches the different categories of semantic information with a stored three-dimensional model of a person-scene to obtain a semantic matching model; and loads and drives the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual image that integrates the person and the scene. The method provided by the embodiment of the present invention utilizes a deep learning algorithm to not only monitor the physical condition and behavioral patterns of the monitored person in real time, but also improves the automation and intelligence of caregiving, eliminating the need for human intervention, reducing caregiving costs, helping caregivers identify potential problems in advance, and providing more scientific and timely decision-making support. Furthermore, the embodiment combines information about the interaction between the person and the scene for semantic reconstruction, improving the accuracy and realism of the reconstruction, and enhancing the efficiency and accuracy of caregiving.
[0078] In one embodiment, the present invention provides an electronic device having the above-mentioned intelligent care system.
[0079] An embodiment of the present invention provides an intelligent nursing system and method, wherein the intelligent nursing system uses a deep learning algorithm to monitor the physical condition and behavior patterns of the monitored party in real time, and combines scene information for semantic analysis to help timely discover and deal with potential problems, thereby improving the automation and intelligence of nursing. In addition, a small amount of semantic information transmitted between the server and the monitoring party is used to reconstruct information about the monitored party and the scene, reducing the risk of privacy exposure of the monitored party and improving the security of data transmission.
[0080] The above-described embodiments merely represent several preferred implementations of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art could make several improvements and substitutions without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be based on the scope of protection of the claims.
Claims
1. An intelligent nursing system, characterized in that: Applied to an intelligent monitoring system, the intelligent monitoring system includes a data analysis server and multiple monitoring edge terminals and multiple visual perception terminals that establish communication connections with the data analysis server, wherein the visual perception terminal and the data analysis server are located in the same local area network, the visual perception terminal includes a monocular camera, and the system includes: The visual perception terminal is used to capture the original nursing video stream through the monocular camera and upload it to the data analysis server; The data analysis server is configured to perform semantic extraction on the original image frames in the original nursing video stream using a deep learning algorithm to obtain semantic information of different categories, and transmit the semantic information to the monitoring edge end via a preset network transmission protocol; wherein the semantic information includes human semantic information, scene semantic information, and human-scene interaction status information between the human body and the scene; The monitoring edge is used to match different categories of semantic information with its stored three-dimensional model of people and scenes to obtain a semantic matching model, and load and drive the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual image that integrates the human body and the scene; The data analysis server is specifically configured to extract human semantic information and scene semantic information from the original image frames of the original nursing video stream using a deep learning algorithm, and analyze the human semantic information and scene semantic information using a pre-built human interaction knowledge graph to obtain human-scene interaction state information between the human body and the scene; The human body semantic information includes human body behavior state information and human body position information. The human body semantic information is extracted from the original image frame of the original nursing video stream using a deep learning algorithm, specifically: Using a pre-trained target detection model to identify the original image frame to obtain a human target detection frame; Inputting the human target detection frame and the original image frame into an improved posture estimation model to obtain human skeleton point coordinate information; the improved posture estimation model is a posture estimation model embedded with a posture structured expression algorithm; Inputting the human skeleton point coordinate information and the original image frame into a spatiotemporal graph convolutional network model to obtain character behavior state information; The preset target skeleton point coordinate information is obtained from the human body skeleton point coordinate information, and the target skeleton point coordinate information is converted into the human body three-dimensional coordinates in the world coordinate system by using camera calibration to obtain the human body position information.
2. The intelligent nursing system according to claim 1, wherein: The semantic information also includes event information, and the event information includes normal event information and abnormal event information. The data analysis server is further configured to: The event information is transmitted to the monitoring edge end through a preset network transmission protocol, so that the monitoring edge end monitors the event information and generates alarm information according to the abnormal event information when abnormal event information is detected.
3. The intelligent nursing system according to claim 1, wherein: The 3D reconstructed virtual image is an interactive 3D reconstructed virtual image, and the monitoring edge end is further used to: When a model interaction trigger event is detected, in response to the model interaction trigger event, model identification data corresponding to the model interaction trigger event is obtained, and human-scene three-dimensional model information corresponding to the model identification data is retrieved and displayed according to the model identification data.
4. The intelligent nursing system according to claim 1, wherein: The human-scene 3D model includes a human body 3D model and a scene 3D model; The human-scene three-dimensional model is stored in a three-dimensional model library at the monitoring edge end. The three-dimensional model library is stored in a relational database management system in the form of a server path. The relational database management system includes an event management unit for storing event information and a model management unit for storing all human-scene three-dimensional model information.
5. The intelligent nursing system according to claim 1, wherein: The preset network transmission protocol includes the User Datagram Protocol.
6. An intelligent nursing method, characterized in that: The method is applied to an intelligent monitoring system, the intelligent monitoring system including a data analysis server and multiple monitoring edge terminals and multiple visual perception terminals that establish communication connections with the data analysis server, wherein the visual perception terminals and the data analysis server are located in the same local area network, and the visual perception terminals include a monocular camera. The method includes the following steps: Capturing and uploading the original nursing video stream through the monocular camera; Using a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain different categories of semantic information; wherein the semantic information includes human semantic information, scene semantic information, and human-scene interaction state information between the human body and the scene; Matching different categories of semantic information with the stored three-dimensional model of human-scene to obtain a semantic matching model, and loading and driving the semantic matching model corresponding to each category of semantic information to generate a 3D reconstructed virtual image that integrates the human body and the scene; The method of using a deep learning algorithm to perform semantic extraction on the original image frames in the original nursing video stream to obtain semantic information of different categories specifically includes: extracting human semantic information and scene semantic information from the original image frames in the original nursing video stream using a deep learning algorithm, and analyzing the human semantic information and scene semantic information using a pre-built human interaction knowledge graph to obtain human-scene interaction state information between the human body and the scene; The human body semantic information includes human body behavior state information and human body position information. The human body semantic information is extracted from the original image frame of the original nursing video stream using a deep learning algorithm, specifically: Using a pre-trained target detection model to identify the original image frame to obtain a human target detection frame; Inputting the human target detection frame and the original image frame into an improved posture estimation model to obtain human skeleton point coordinate information; the improved posture estimation model is a posture estimation model embedded with a posture structured expression algorithm; Inputting the human skeleton point coordinate information and the original image frame into a spatiotemporal graph convolutional network model to obtain character behavior state information; The preset target skeleton point coordinate information is obtained from the human body skeleton point coordinate information, and the target skeleton point coordinate information is converted into the human body three-dimensional coordinates in the world coordinate system by using camera calibration to obtain the human body position information.
7. An electronic device, characterized in that: A smart nursing system according to any one of claims 1 to 5.
Citation Information
Patent Citations
Outdoor monocular synchronous mapping and positioning method fusing scene semantics
CN112734845A
AR portrait photographing method and system based on 3D positioning information
CN113643357A