Clinical patient early warning nursing system and method based on deep learning
Through a deep learning-based clinical patient warning and care system, the video images collected by multi-lens cameras are sorted and stitched, which solves the ghosting problem during multi-image fusion, and ensures the synchronization effect between audio and video through audio processing, achieving higher monitoring accuracy and reliability.
Patent Information
- Application Number
- CN202510184461.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing postoperative surveillance and nursing equipment for critically ill patients is not perfect enough. When taking multiple pictures, multi-lens cameras cause ghosting due to angle problems, making perfect multi-image fusion impossible.
The clinical patient early warning and nursing system based on deep learning is adopted. The video images collected by multiple lenses are sorted and spliced through the image acquisition module and the image adjustment module to avoid ghosting, and the audio acquisition module and the audio adjustment module are processed and synthesized to ensure that the audio sound effect position matches the video picture.
It realizes the perfect fusion of multiple images, avoids ghosting, and ensures the synchronization effect of audio and video, improving the accuracy and reliability of patient monitoring.
Smart Images

Figure CN120108786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to patient care technology, and in particular to a clinical patient early warning care system and method based on deep learning. Background Art
[0002] Medical equipment refers to instruments, equipment, appliances, materials or other items used on the human body alone or in combination, including the required software. Medical equipment is the most basic element of medical, scientific research, teaching, institutional and clinical discipline work, including professional medical equipment and home medical equipment. Medical equipment is the basic condition for continuously improving the level of medical science and technology, and is also an important symbol of the degree of modernization. Medical equipment has become an important field of modern medicine. The development of medical care depends to a large extent on the development of instruments. Even in the development of the medical industry, its breakthrough of bottlenecks has played a decisive role. Medical equipment refers to instruments, equipment, appliances, materials or other items used on the human body alone or in combination, including the required software. The therapeutic effect on the human body surface and body is not obtained by means of pharmacology, immunology or metabolism, but medical device products play a certain auxiliary role. During use, it is intended to achieve the following intended purposes: prevention, diagnosis, treatment, monitoring, and relief of diseases; diagnosis, treatment, monitoring, relief, and compensation of injuries or disabilities; research, substitution, and regulation of anatomical or physiological processes; and pregnancy control.
[0003] The existing postoperative monitoring and nursing equipment for critically ill patients is not perfect. Multi-lens nursing monitoring cameras are usually used to comprehensively monitor patients. However, during the use of multi-lens nursing monitoring cameras, due to the shooting angle problem of the multi-lens camera, the multiple pictures taken cannot be perfectly integrated, resulting in ghosting when multiple images are integrated. Summary of the invention
[0004] The purpose of the present invention is to provide a clinical patient early warning nursing system and method based on deep learning to solve the above-mentioned deficiencies in the prior art.
[0005] In order to achieve the above object, the present invention provides the following technical solution: a clinical patient early warning nursing method based on deep learning, comprising the following working steps:
[0006] S1, collecting pictures of different video images captured by multiple lenses at the same time, and sorting the multiple different pictures;
[0007] S2, based on the sorting result in step S1, stitching two adjacent pictures;
[0008] S3, continuously processing the video images captured by adjacent lenses based on the splicing scheme of the two adjacent screen images in step S2;
[0009] S4, collecting the picture audio collected when multiple lenses capture video images, and performing audio processing on the picture audio, wherein the audio processing includes audio denoising, audio keyword recognition, and audio volume adjustment;
[0010] S5, performing picture and audio matching on the video image processed in step S3 and the multiple pictures and sounds pre-processed in step S4 based on a data matching algorithm, synthesizing the pictures and audios with matching results, and performing audio adjustment on the pictures and audios without matching results;
[0011] S6, performing audio adjustment on the picture and sound for which no matching result exists in step S5 based on the audio waveform adjustment algorithm, and synthesizing the adjusted audio with the video image, wherein the specific standard of the audio waveform adjustment is to adjust the sound effect position of the audio in the video picture by adjusting the audio waveform;
[0012] S7, based on the video detection algorithm, simultaneously detects the lens image and the audio playback effect to ensure that the audio sound effect position matches the video image when the video is played.
[0013] Furthermore, the step S1 of sorting the multiple different video images specifically includes the following steps:
[0014] A1, taking panoramic images by a panoramic camera;
[0015] A2, dividing the panoramic video captured by the panoramic camera according to the horizontal relative angle and the vertical relative angle of the panoramic camera by a panoramic video angle division module, to obtain the field of view angle corresponding to each position of the panoramic video;
[0016] A3, using a picture recognition algorithm to respectively identify the surrounding image information of different pictures collected by multiple lenses;
[0017] A4, performing image matching on the surrounding image information of the picture identified in step A3 and the panoramic image captured in step A1 based on an image matching algorithm, and numbering the matched picture according to the field of view angle obtained in step A2;
[0018] A5, in step A4, the pictures with different horizontal viewing angles are numbered from left to right in the first level, and the pictures with the same horizontal viewing angle but different vertical viewing angles are numbered from top to bottom in the second level.
[0019] Furthermore, the image stitching of two adjacent screen images in step S2 specifically includes the following steps:
[0020] B1, performing adjacent end feature segmentation based on the two pictures with adjacent primary labels in step A5 and the two pictures with the same primary label but adjacent secondary labels in step A5;
[0021] B2, performing cross feature optimization on two adjacent pictures after feature segmentation based on a graph optimization algorithm, wherein the cross feature is a shot picture in which both adjacent pictures exist;
[0022] B3, based on the image synthesis algorithm, synthesize the picture in which the adjacent end features are segmented and there is no cross feature in step B1 and the cross picture optimized in step B2.
[0023] Furthermore, the feature segmentation in step B1 specifically includes the following steps:
[0024] C1, in order to obtain the detailed features of different spatial dimensions of the picture to improve the accuracy of feature recognition, the feature map is segmented along the vertical and horizontal dimensions. The two independent feature maps output by the backbone network are named FH and FW, namely the vertical branch and the horizontal branch, and then they are segmented into several sub-regions in the vertical and horizontal directions, as shown in the following formula:
[0025]
[0026] Where N represents the number of samples in the data set; f h (i, p1), f w (i, p2) represents the vertical and horizontal block feature maps of the i-th image respectively; p1 and p2 represent the number of segmentations in the vertical and horizontal directions respectively;
[0027] C2, assuming that the feature is C, adopts a multi-granularity segmentation method, and the specific expression is shown in the following formula:
[0028] F C ={f c (i, 1), f c (i, 2), ..., f c (i, p 3 )}i=1,2...N,
[0029] Where N represents the number of samples in the feature data set, FC represents the overall feature, and fc(i, p3) represents the feature of the i-th segmented image;
[0030] C3, performing cross-feature optimization on the pictures with the same features as the i-th segmented image in step C2, and storing the pictures with different features from the i-th segmented image in step C2 for subsequent image synthesis;
[0031] C4, performing a preprocessing operation on the image stored in step C3, wherein the preprocessing operation includes denoising, deblurring, and normalization;
[0032] C5, converting the preprocessed image in step C4 into a format that can be processed by the convolutional neural network, and building a convolutional neural network model;
[0033] C6, using a labeled dataset for model training. The dataset may contain multiple images and their corresponding correct or incorrect segmentation results. During training, these images may be input into the model so that it automatically learns how to correctly segment the image into different regions. During the training process, cross-validation and other techniques may be used to avoid overfitting to improve the generalization ability of the model.
[0034] C7, after completing the training in step C6, use the unlabeled test data set to evaluate and optimize the model;
[0035] C8, outputting the image segmentation result according to the optimal performance of the model in step C7, wherein the image segmentation criterion in the image segmentation result is whether the image content involves the patient's privacy life, if the image segmentation result involves the patient's privacy life, the image is segmented and deleted, if the image segmentation result does not involve the patient's privacy life, the image is segmented and retained;
[0036] C9, adjust the model parameters to obtain the best performance. Specifically, the model can be tuned by adjusting the size of the convolution kernel, step size, regularization parameters and other hyperparameters, as well as by changing the network structure, adding new modules, etc., so that the model has higher segmentation accuracy on the test data set;
[0037] C10 deploys the trained model in real time to achieve real-time image segmentation. In actual application scenarios, the real-time image can be preprocessed and then segmented using the trained model to obtain the required image segmentation results. At the same time, the original image can be further processed according to the segmentation results, such as cropping, scaling, flipping, etc.
[0038] Furthermore, the audio processing of the picture audio in step S4 specifically includes the following steps:
[0039] D1, adjust the audio volume of the collected picture and audio to ensure that the processed picture and audio are of uniform volume;
[0040] D2, performing audio denoising based on the picture audio whose volume has been adjusted in step D1, specifically comprising the following steps:
[0041] D21, for the noisy frequency signal, a section containing only the noise signal is reserved;
[0042] D22, after the noisy frequency signal is preliminarily denoised by the improved multi-window spectrum subtraction method, it is also necessary to adjust the noise signal collected according to the spectrum subtraction and gain parameters transmitted in the first stage of denoising, such as reducing the signal amplitude;
[0043] D23, then the adjusted noise signal is used as the actual reference signal and input into the filter together with the audio signal after the first stage of denoising, so as to complete the overall denoising process and obtain a relatively pure audio signal;
[0044] D24, finally, the audio signal after denoising can be intercepted and processed according to the time interval size of the interval containing only noise signals reserved in the previous stage to obtain a pure audio signal;
[0045] D3, extracting audio keywords from the denoised pure audio signal, specifically including the following steps:
[0046] D31, feature extraction is to extract the speech features of the collected audio signal. This step can effectively enhance the key features of the audio signal and reduce the computational dimension of keyword recognition;
[0047] D32, the neural network part is to classify the key information obtained after feature extraction, so as to obtain the result after keyword recognition in the audio signal;
[0048] D4, performing subsequent processing based on the noise-free frequency keywords generated in steps D3 and D2, wherein the subsequent processing includes synthesizing the audio and video images and adjusting the audio based on an audio waveform adjustment algorithm.
[0049] A clinical patient early warning nursing system based on deep learning, including an image acquisition module, an image adjustment module, an audio acquisition module, an audio adjustment module, and a video detection module:
[0050] The image acquisition module is used to collect different video images acquired by multiple lenses, and according to the time when the multiple lenses acquired the images, the different video images acquired by the multiple lenses at the same time are uniformly processed, and the screen images are sorted;
[0051] The image adjustment module is used to receive different video images collected by multiple lenses at the same time and sent by the image acquisition module, and to splice two adjacent pictures;
[0052] The audio acquisition module is used to collect the audio of the picture collected when multiple lenses capture video images, and perform audio processing on the audio of the picture, wherein the audio processing includes audio denoising, audio keyword recognition and audio volume adjustment;
[0053] The audio adjustment module is used to match the processed video image with the processed multiple pictures and sounds by using a data matching algorithm, and synthesize the pictures and audios for which matching results exist; for pictures and sounds for which matching results do not exist, perform audio adjustment on the sounds based on an audio waveform adjustment algorithm, and synthesize the adjusted audio with the video image;
[0054] The video detection module detects the lens image and audio playback effect simultaneously based on the video detection algorithm;
[0055] The early warning broadcast module is used to detect the video images that have passed the detection by the video detection module based on the early warning broadcast algorithm, and to issue an early warning broadcast when the patient has abnormal behavior or abnormal state.
[0056] Compared with the prior art, the clinical patient early warning nursing system and method based on deep learning provided by the present invention, through the image acquisition module and the image adjustment module, when identifying the video images captured by multiple lenses, by identifying the range of the field of view angle shot by the corresponding lens, the pictures captured by multiple lenses are sorted according to the field of view angle, so as to ensure the integrity of the multiple pictures, and at the same time, feature segmentation can be used to avoid ghosting when multiple images are synthesized during picture splicing, and at the same time, the audio acquisition module and the audio adjustment module are set to adjust the audio effect position in the video picture by adjusting the audio waveform, so as to ensure that the audio sound effect position matches the video picture when the video is played. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0058] Figure 1 A schematic diagram of the overall method flow provided by an embodiment of the present invention;
[0059] Figure 2 A schematic diagram of the overall structure provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0061] See also Figure 1 A clinical patient early warning nursing method based on deep learning, a clinical patient early warning nursing method based on deep learning, including the following working steps:
[0062] S1, collecting pictures of different video images captured by multiple lenses at the same time, and sorting the multiple different pictures;
[0063] S2, based on the sorting result in step S1, stitching two adjacent pictures;
[0064] S3, continuously processing the video images captured by adjacent lenses based on the splicing scheme of the two adjacent screen images in step S2;
[0065] S4, collecting the audio of the picture collected when multiple lenses collect video images, and performing audio processing on the audio of the picture, the audio processing including audio denoising, audio keyword recognition, and audio volume adjustment;
[0066] S5, performing picture and audio matching on the video image processed in step S3 and the multiple pictures and sounds pre-processed in step S4 based on a data matching algorithm, synthesizing the pictures and audios with matching results, and performing audio adjustment on the pictures and audios without matching results;
[0067] S6, performing audio adjustment on the picture and sound for which no matching result exists in step S5 based on the audio waveform adjustment algorithm, and synthesizing the adjusted audio with the video image, wherein the specific standard of the audio waveform adjustment is to adjust the sound effect position of the audio in the video picture by adjusting the audio waveform;
[0068] S7, based on the video detection algorithm, simultaneously detects the lens image and the audio playback effect to ensure that the audio sound effect position matches the video image when the video is played;
[0069] S8, based on the early warning broadcast module, is used to detect the qualified video images detected in step S7 based on the early warning broadcast algorithm, and to issue an early warning broadcast when the patient has abnormal behavior or abnormal state.
[0070] By setting up the image acquisition module and the image adjustment module in this way, when identifying video images captured by multiple lenses, the range of the field of view angles shot by the corresponding lenses is identified, and the pictures captured by multiple lenses are sorted according to the field of view angles, thereby ensuring the integrity of the multiple pictures. At the same time, feature segmentation can be used to avoid ghosting when multiple images are synthesized during image stitching. At the same time, the audio acquisition module and the audio adjustment module are set to adjust the audio effect position in the video screen by adjusting the audio waveform, thereby ensuring that the audio sound effect position matches the video screen during video playback. The patient image can be segmented for privacy through feature segmentation of the image to avoid leakage of patient privacy during monitoring. At the same time, the early warning broadcast module can be set to detect video pictures that have passed the video detection module detection based on the early warning broadcast algorithm, thereby realizing early warning broadcasts when patients have abnormal behaviors or abnormal states.
[0071] Sorting the multiple different video images in step S1 specifically includes the following steps:
[0072] A1: First of all, as a key image acquisition device, the panoramic camera, with its unique perspective and wide field of view, can capture all-round panoramic images and record the surrounding environment in a complete way, providing rich materials for subsequent processing and analysis.
[0073] A2: Next, the panoramic video angle division module plays an important role in dividing the panoramic video captured by the panoramic camera in detail. According to the two key dimensions of the horizontal relative angle and the vertical relative angle of the panoramic camera, the entire panoramic video is divided into small areas, so as to obtain the field of view angle corresponding to each position of the panoramic video. These field of view angles are like small windows, corresponding to the viewing angle range of different positions in the panoramic video, laying the foundation for subsequent picture recognition and matching work.
[0074] A3: Then, the image recognition algorithm comes into play. It is like a keen image detective, which can carefully observe and analyze different images collected by multiple lenses. Through its powerful recognition ability, it can accurately identify the surrounding image information of these images, just like capturing the edge details of the image from all directions, providing comprehensive information for subsequent image matching.
[0075] A4: On this basis, the image matching algorithm comes into play. The surrounding image information of the picture identified in step A3 is cleverly matched with the panoramic image taken in step A1. Through complex algorithm calculations, the matching pictures are associated with the corresponding positions in the panoramic image, just like finding the missing puzzle pieces. Then, the matched pictures are numbered in order according to the field of view angle obtained in step A2, so that each picture can find its own position and belonging in the panoramic image.
[0076] A5: Finally, based on step A4, for the pictures with different horizontal viewing angles, first-level numbering is performed from left to right, just like labeling pictures with different viewing angles on a horizontal line, clearly marking their order in the horizontal direction. For the pictures with the same horizontal viewing angle but different vertical viewing angles in step A4, second-level numbering is performed from top to bottom, just like subdividing and numbering them in the vertical direction on the same horizontal line. In this way, pictures with different viewing angles can be managed and identified more meticulously, providing more accurate and detailed information for subsequent applications and analysis.
[0077] The step S2 of stitching two adjacent pictures comprises the following steps:
[0078] B1, performing adjacent end feature segmentation based on the two pictures with adjacent primary labels in step A5 and the two pictures with the same primary label but adjacent secondary labels in step A5;
[0079] B2, based on the graph optimization algorithm, the cross-feature optimization of two adjacent images after feature segmentation is studied to improve the accuracy and efficiency of image processing. Cross-features refer to the shooting picture features that exist in two adjacent images. Through the graph optimization algorithm, the similar information of the two images can be effectively integrated, thereby achieving the purpose of enhancing feature extraction and segmentation effects;
[0080] In the specific implementation, the graph optimization algorithm is first used to detect and extract feature points from two adjacent images. Then, by analyzing the similarities and differences of the cross-features, a series of optimization strategies are formulated to achieve accurate matching and enhancement of features. This process not only improves the overall quality of the image, but also lays a good foundation for subsequent image analysis, recognition, and classification applications.
[0081] In summary, by optimizing the cross-features of two adjacent images, the effect and application scope of graphics processing can be effectively improved, providing new ideas and methods for the further development of computer vision.
[0082] B3, specifically, in step B1, firstly, the adjacent end images are subjected to detailed feature segmentation to ensure that the important features in each image can be accurately identified and extracted. Since there are no cross features in these images, they provide independent and important visual information. On this basis, step B2 refines the cross feature images through an image optimization algorithm to ensure that the generated cross feature images can effectively reflect the key information coexisting in the adjacent images.
[0083] Combining the results of step B1 and step B2, an image synthesis algorithm is used to fuse the adjacent end images that do not contain cross features with the optimized cross feature image to generate a new synthetic image. This synthesis process not only considers the visual attributes of the image such as color and brightness, but also takes into account the coherence and consistency of the image content. Through this method, the final synthetic image will have higher visual quality and can effectively show the relationship and characteristics between adjacent images.
[0084] In summary, the synthesis method based on the image synthesis algorithm proposed in the present invention can generate high-quality synthetic images based on feature segmentation and cross-feature optimization. It is widely used in image editing, visual effect enhancement, computer vision and other fields, and provides strong support for the development of related technologies.
[0085] The feature segmentation in step B1 specifically includes the following steps:
[0086] C1, in order to obtain the detailed features of different spatial dimensions of the picture to improve the accuracy of feature recognition, the feature map is segmented along the vertical and horizontal dimensions. The two independent feature maps output by the backbone network are named FH and FW, namely the vertical branch and the horizontal branch, and then they are segmented into several sub-regions in the vertical and horizontal directions, as shown in the following formula:
[0087]
[0088] Where N represents the number of samples in the data set; f h (i, p1), f w (i, p2) represents the vertical and horizontal block feature maps of the i-th image respectively; p1 and p2 represent the number of segmentations in the vertical and horizontal directions respectively;
[0089] C2, assuming that the feature is C, adopts a multi-granularity segmentation method, and the specific expression is shown in the following formula:
[0090] F C ={f c (i, 1), f c (i, 2), ..., f c (i, p 3 )}i=1,2...N,
[0091] Where N represents the number of samples in the feature data set, FC represents the overall feature, and fc(i, p3) represents the feature of the i-th segmented image;
[0092] C3, for images with the same features as the i-th segmented image in step C2, a cross-feature optimization algorithm will be used for comprehensive processing. This algorithm aims to identify and extract common features in images, thereby achieving feature sharing and enhancement between different images. In this process, relying on image feature matching and fusion technology, images with the same features are accurately interactively processed to ensure that each similar feature is highlighted. Through this optimization, not only the recognizability of the features is improved, but also greater convenience is created for image synthesis, effectively improving the quality and effect of image processing;
[0093] On the other hand, for images with different features from the i-th segmented image in step C2, they are effectively stored. These images with different features provide rich visual information, and retaining them will provide more possibilities and creative space for subsequent image synthesis. In the storage process, efficient data structures and processing technologies will be used to ensure fast access and dynamic update of data, thereby improving the flexibility and efficiency of subsequent image editing and synthesis;
[0094] In summary, the present invention forms an intelligent and efficient image processing framework by cross-feature optimization of images with the same features and efficient storage of images with different features, laying a solid foundation for the subsequent image synthesis process. This method not only improves the accuracy of image processing, but also provides broader possibilities for image generation and editing applications, and promotes the further development of computer vision and image synthesis technology;
[0095] C4, performing a preprocessing operation on the image stored in step C3, the preprocessing operation includes denoising, deblurring, and normalization;
[0096] C5, converting the preprocessed image in step C4 into a format that can be processed by the convolutional neural network, and building a convolutional neural network model;
[0097] C6, use labeled datasets for model training. The dataset can contain multiple images and their corresponding correct or incorrect segmentation results. During training, these images can be input into the model to automatically learn how to correctly segment the image into different regions. During the training process, cross-validation and other techniques can be used to avoid overfitting to improve the generalization ability of the model.
[0098] C7, after completing the training in step C6, use the unlabeled test data set to evaluate and optimize the model;
[0099] C8, outputting the image segmentation result according to the optimal performance of the model in step C7, wherein the image segmentation criterion in the image segmentation result is whether the image content involves the patient's privacy life, if the image segmentation result involves the patient's privacy life, the image is segmented and deleted, if the image segmentation result does not involve the patient's privacy life, the image is segmented and retained;
[0100] C9, adjust the model parameters to obtain the best performance. Specifically, the optimization process mainly includes the following aspects:
[0101] 1. Convolution kernel parameter adjustment: By accurately adjusting the size of the convolution kernel, features at different scales can be effectively captured. Larger convolution kernels help extract higher-level abstract features, while smaller convolution kernels focus on detailed features. According to specific task requirements, you can choose an appropriate convolution kernel size to improve the model's sensitivity to features;
[0102] 2. Step size and padding strategy: The choice of step size directly affects the receptive field of the model and the size of the output feature map. Reasonable step size setting can balance the computational complexity and information integrity. Too large a step size may lead to the loss of feature information, while too small a step size may increase the amount of computation. In addition, the choice of padding strategy can also determine the degree to which the model retains boundary information, further optimizing the segmentation effect;
[0103] 3. Regularization parameters: By adjusting the regularization parameters (such as L1 or L2 regularization), the overfitting problem of the model can be effectively prevented and the performance of the model on unseen data can be improved. This process requires continuous fine-tuning based on the performance of the model on the validation data set in order to obtain the optimal parameter settings;
[0104] 4. Improvement of network structure: Based on the performance of the current model, you can consider improving the network structure. For example, add new convolutional layers or pooling layers to increase the depth and expression ability of the model; or use residual connection, dense connection and other technologies to improve the efficiency of information flow and enhance the stability of gradient transfer during training;
[0105] 5. Modular design: Introducing new modules, such as attention mechanisms or adaptive enhancement modules, can help the model pay more attention to important feature areas and improve the model's ability to understand complex scenes. Such modules can dynamically adjust feature weights to achieve better results in different tasks;
[0106] 6. Hyperparameter optimization: Use cross-validation, grid search, Bayesian optimization and other techniques to systematically tune the model's hyperparameters. By searching in the preset parameter space, the optimal hyperparameter combination can be found to maximize model performance;
[0107] C10 deploys the trained model in real time to achieve real-time image segmentation. In actual application scenarios, the real-time image can be preprocessed and then segmented using the trained model to obtain the required image segmentation results. At the same time, the original image can be further processed according to the segmentation results, such as cropping, scaling, flipping, etc.
[0108] The audio processing of the picture audio in step S4 specifically includes the following steps:
[0109] D1, adjust the audio volume of the collected picture and audio to ensure that the processed picture and audio are of uniform volume;
[0110] D2, performing audio denoising based on the picture audio whose volume has been adjusted in step D1, specifically comprising the following steps:
[0111] D21, for the noisy frequency signal, a section containing only the noise signal is reserved;
[0112] D22, after the noisy frequency signal is preliminarily denoised by the improved multi-window spectrum subtraction method, it is also necessary to adjust the noise signal collected according to the spectrum subtraction and gain parameters transmitted in the first stage of denoising, such as reducing the signal amplitude;
[0113] D23, then the adjusted noise signal is used as the actual reference signal and input into the filter together with the audio signal after the first stage of denoising, so as to complete the overall denoising process and obtain a relatively pure audio signal;
[0114] D24, finally, the audio signal after denoising can be intercepted and processed according to the time interval size of the interval containing only noise signals reserved in the previous stage to obtain a pure audio signal;
[0115] D3, extracting audio keywords from the denoised pure audio signal, specifically including the following steps:
[0116] D31, feature extraction is to extract the speech features of the collected audio signal. This step can effectively enhance the key features of the audio signal and reduce the computational dimension of keyword recognition;
[0117] D32, the neural network part is to classify the key information obtained after feature extraction, so as to obtain the result of keyword recognition in the audio signal. The neural network part is responsible for effectively classifying the key information obtained through the feature extraction stage to achieve accurate recognition of keywords in the audio signal. In this process, the neural network identifies the features in different audio signals through multiple levels of nonlinear transformation, and finally outputs the results of keyword recognition;
[0118] Specifically, the neural network architecture includes:
[0119] Feature extraction layer: Before inputting the audio signal, the feature information in the audio signal is first extracted through preprocessing techniques (such as Fourier transform or Mel-frequency cepstral coefficient MFCC extraction). These features make the data easier to process in the subsequent classification process, thereby improving the performance of the model;
[0120] Convolutional layer: Using the convolutional layer in the convolutional neural network (CNN) can effectively capture local features in the audio signal. Through the convolution operation, the model can extract the time domain and frequency domain features in the audio signal and enhance the information expression ability of the important part of the speech signal. Appropriate convolution kernel design can help the model focus on key information;
[0121] Pooling layer: The pooling layer is used to downsample the feature map, reduce the amount of calculation and prevent overfitting. The maximum pooling and average pooling operations can effectively retain important features while reducing the complexity of network parameters and enhancing the robustness of the model;
[0122] Fully connected layer: In the final stage of the neural network, the features are further processed through the fully connected layer. Here, each neuron is connected to all neurons in the previous layer to comprehensively consider all the extracted feature information. Through activation functions (such as ReLU or softmax), the model can be mapped to different keyword categories;
[0123] Classification output layer: The final classification output layer will map the features in the audio signal through the weights generated by training to generate the final result of keyword recognition. At this time, the model will output the probability distribution of the corresponding keywords and be able to determine the most likely keywords in the audio signal;
[0124] Training and optimization: The training process of the neural network uses the back propagation algorithm, combined with the cross entropy loss function, to optimize the parameters of the model. Through iterative training, the model can achieve a balance between recognition accuracy and loss value, and gradually improve the generalization ability of unseen data;
[0125] D4, performing subsequent processing based on the noise-free frequency keywords generated in steps D3 and D2, the subsequent processing including synthesizing the audio and video images and adjusting the audio based on an audio waveform adjustment algorithm.
[0126] See also Figure 2 , a clinical patient early warning nursing system based on deep learning, including an image acquisition module, an image adjustment module, an audio acquisition module, an audio adjustment module, and a video detection module:
[0127] The image acquisition module is used to collect different video images collected by multiple lenses, and according to the time when the multiple lenses collected the images, the different video images collected by multiple lenses at the same time are uniformly processed and the pictures are sorted;
[0128] The image adjustment module is used to receive different video images collected by multiple lenses at the same time and sent by the image acquisition module, and to stitch two adjacent pictures;
[0129] The audio acquisition module is used to collect the audio of the picture collected when multiple lenses capture video images, and perform audio processing on the picture audio. The audio processing includes audio denoising, audio keyword recognition and audio volume adjustment.
[0130] The audio adjustment module is used for performing picture and audio matching between the processed video image and the processed multiple pictures and sounds by using a data matching algorithm, and synthesizing the pictures and audios for which matching results exist; for pictures and sounds for which matching results do not exist, performing audio adjustment on the sounds based on an audio waveform adjustment algorithm, and synthesizing the adjusted audio with the video image;
[0131] The audio waveform adjustment includes the following steps:
[0132] Audio waveform analysis: First, perform waveform analysis on the unmatched audio signal. By observing the audio waveform, we can identify the frequency components, amplitude, duration, and change characteristics of the audio signal. Using time domain and frequency domain analysis, we can reveal which parts of the audio and video content are inconsistent, thus providing a basis for subsequent adjustments;
[0133] Sound effect position adjustment: Through the audio waveform adjustment algorithm, we can change the position of the sound effect within the picture. This process involves repositioning or remixing specific frequencies or various sound objects (such as dialogue, music, ambient sound, etc.) in the audio signal through audio positioning technology to adapt to the spatial layout of the video screen. For example, if a character on the left side of the video screen is speaking, and the dialogue part of the original audio is on the right, the sound effect position of the dialogue can be adjusted to make it more naturally move with the character in the picture, thereby improving the consistency of vision and hearing;
[0134] Intelligent adjustment of audio features: The audio waveform adjustment algorithm will also make intelligent adjustments to certain audio elements based on the frequency and pitch characteristics of the audio. For example, the bass frequency of the background music can be enhanced to ensure that it does not cover up the dialogue and enhance the emotional expression; or through dynamic range compression, reverberation and other technologies, the application effect of the audio in different environments can be ensured. At the same time, the rhythm and duration of the audio are corrected to ensure that it is consistent with the rhythm of the corresponding video screen;
[0135] Synthesis and output: After the audio adjustment is completed, the adjusted audio and video images are synthesized. At this stage, it is necessary to ensure that the audio and video are accurately aligned on the timeline to avoid lag or advance. During the synthesis process, you can consider using advanced audio encoding technology to maintain sound quality and reduce file size, ensuring that the final multimedia work can be played smoothly on various devices;
[0136] Final inspection and optimization: After completing the synthesis of audio and video, the final multimedia product needs to be quality checked. This includes playing and observing the synthesis results to ensure that the sound effects and dynamic performance in the picture are consistent, and the clarity and positioning of the audio should also reach a satisfactory level. If unnatural places are found, secondary adjustments will be made based on the feedback of the results to ensure that the quality of the final output meets professional standards.
[0137] The video detection module detects the lens image and audio playback effects simultaneously based on the video detection algorithm;
[0138] The early warning broadcast module is used to detect the video images that have passed the video detection module based on the early warning broadcast algorithm, and to issue early warning broadcasts when the patient has abnormal behavior or abnormal state.
[0139] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A clinical patient early warning nursing method based on deep learning, characterized in that: The following steps are included: S1, collecting pictures of different video images captured by multiple lenses at the same time, and sorting the multiple different pictures; S2, based on the sorting result in step S1, stitching two adjacent screen images, wherein stitching two adjacent screen images specifically includes the following steps: B1, performing adjacent end feature segmentation based on the two pictures with adjacent primary labels in step A5 and the two pictures with the same primary label but adjacent secondary labels in step A5; B2, performing cross feature optimization on two adjacent pictures after feature segmentation based on a graph optimization algorithm, wherein the cross feature is a shot picture in which both adjacent pictures exist; B3, based on an image synthesis algorithm, synthesize the picture image in which the adjacent end features are segmented and there is no cross feature in step B1 and the cross image optimized in step B2; S3, continuously processing the video images captured by adjacent lenses based on the splicing scheme of the two adjacent screen images in step S2; S4, collecting the picture audio collected when multiple lenses capture video images, and performing audio processing on the picture audio, wherein the audio processing includes audio denoising, audio keyword recognition, and audio volume adjustment; S5, performing picture and audio matching on the video image processed in step S3 and the multiple pictures and sounds pre-processed in step S4 based on a data matching algorithm, synthesizing the pictures and audios with matching results, and performing audio adjustment on the pictures and audios without matching results; S6, performing audio adjustment on the picture and sound for which no matching result exists in step S5 based on the audio waveform adjustment algorithm, and synthesizing the adjusted audio with the video image, wherein the audio waveform adjustment changes the sound effect position in the video picture by adjusting the audio waveform; S7, based on the video detection algorithm, simultaneously detects the lens image and the audio playback effect to ensure that the audio sound effect position matches the video image when the video is played; S8, based on the early warning broadcast module, is used to detect the qualified video images detected in step S7 based on the early warning broadcast algorithm, and to issue an early warning broadcast when the patient has abnormal behavior or abnormal state.
2. A clinical patient early warning nursing method based on deep learning according to claim 1, characterized in that: The adjacent end feature segmentation specifically comprises the following steps: C1, in order to obtain the detailed features of different spatial dimensions of the picture to improve the accuracy of feature recognition, the feature map is segmented along the vertical and horizontal dimensions. The two independent feature maps output by the backbone network are named FH and FW, namely the vertical branch and the horizontal branch, and then they are segmented into several sub-regions in the vertical and horizontal directions, as shown in the following formula: Where N represents the number of samples in the data set; f h (i, p1), f w (i, p2) represents the vertical and horizontal block feature maps of the i-th image respectively; p1 and p2 represent the number of segmentations in the vertical and horizontal directions respectively; C2, assuming that the feature is C, adopts a multi-granularity segmentation method, and the specific expression is shown in the following formula: F C ={f c (i,1),f c (i,2),...,f c (i,p3)}i=1,2...N, Where N represents the number of samples in the feature data set, FC represents the overall feature, and fc(i, p3) represents the feature of the i-th segmented image; C3, performing cross-feature optimization on the pictures with the same features as the i-th segmented image in step C2, and storing the pictures with different features from the i-th segmented image in step C2 for subsequent image synthesis; C4, performing a preprocessing operation on the image stored in step C3, wherein the preprocessing operation includes denoising, deblurring, and normalization; C5, converting the preprocessed image in step C4 into a format that can be processed by the convolutional neural network, and building a convolutional neural network model; C6, using a labeled dataset for model training. The dataset may contain multiple images and their corresponding correct or incorrect segmentation results. During training, these images may be input into the model so that it automatically learns how to correctly segment the image into different regions. During the training process, cross-validation and other techniques may be used to avoid overfitting to improve the generalization ability of the model. C7, after completing the training in step C6, use the unlabeled test data set to evaluate and optimize the model; C8, outputting the image segmentation result according to the optimal performance of the model in step C7, wherein the image segmentation criterion in the image segmentation result is whether the image content involves the patient's privacy life, if the image segmentation result involves the patient's privacy life, the image is segmented and deleted, if the image segmentation result does not involve the patient's privacy life, the image is segmented and retained; C9, adjust model parameters to obtain optimal performance; C10 deploys the trained model in real time to achieve real-time image segmentation. In actual application scenarios, the real-time image can be preprocessed and then segmented using the trained model to obtain the desired image segmentation result.
3. A clinical patient early warning nursing method based on deep learning according to claim 1, characterized in that: The step S1 of sorting the multiple different video images specifically includes the following steps: A1, taking panoramic images by a panoramic camera; A2, dividing the panoramic video captured by the panoramic camera according to the horizontal relative angle and the vertical relative angle of the panoramic camera by a panoramic video angle division module, to obtain the field of view angle corresponding to each position of the panoramic video; A3, using a picture recognition algorithm to respectively identify the surrounding image information of different pictures collected by multiple lenses; A4, performing image matching on the surrounding image information of the picture identified in step A3 and the panoramic image captured in step A1 based on an image matching algorithm, and numbering the matched picture according to the field of view angle obtained in step A2; A5, in step A4, the pictures with different horizontal viewing angles are numbered from left to right in the first level, and the pictures with the same horizontal viewing angle but different vertical viewing angles are numbered from top to bottom in the second level.
4. A clinical patient early warning nursing method based on deep learning according to claim 1, characterized in that: The audio processing of the picture audio in step S4 specifically includes the following steps: D1, adjust the audio volume of the collected picture and audio to ensure that the processed picture and audio are of uniform volume; D2, performing audio denoising based on the picture audio whose volume has been adjusted in step D1, specifically comprising the following steps: D21, for the noisy frequency signal, a section containing only the noise signal is reserved; D22, after the noisy frequency signal is preliminarily denoised by the improved multi-window spectrum subtraction method, it is also necessary to adjust the noise signal collected according to the spectrum subtraction and gain parameters transmitted in the first stage of denoising, such as reducing the signal amplitude; D23, then the adjusted noise signal is used as the actual reference signal and input into the filter together with the audio signal after the first stage of denoising, so as to complete the overall denoising process and obtain a relatively pure audio signal; D24, finally, the audio signal after denoising can be intercepted and processed according to the time interval size of the interval containing only noise signals reserved in the previous stage to obtain a pure audio signal; D3, extracting audio keywords from the denoised pure audio signal, specifically including the following steps: D31, feature extraction is to extract the speech features of the collected audio signal. This step can effectively enhance the key features of the audio signal and reduce the computational dimension of keyword recognition; D32, the neural network part is to classify the key information obtained after feature extraction, so as to obtain the result after keyword recognition in the audio signal; D4, performing subsequent processing based on the noise-free frequency keywords generated in steps D3 and D2, wherein the subsequent processing includes synthesizing the audio and video images and adjusting the audio based on an audio waveform adjustment algorithm.
5. A clinical patient early warning nursing system based on deep learning, which is applicable to a clinical patient early warning nursing method based on deep learning according to any one of claims 1 to 4, characterized in that: Including image acquisition module, image adjustment module, audio acquisition module, audio adjustment module, video detection module, and early warning broadcast module: The image acquisition module is used to collect different video images acquired by multiple lenses, and according to the time when the multiple lenses acquired the images, the different video images acquired by the multiple lenses at the same time are uniformly processed, and the screen images are sorted; The image adjustment module is used to receive different video images collected by multiple lenses at the same time and sent by the image acquisition module, and to splice two adjacent pictures; The audio acquisition module is used to collect the audio of the picture collected when multiple lenses capture video images, and perform audio processing on the audio of the picture, wherein the audio processing includes audio denoising, audio keyword recognition and audio volume adjustment; The audio adjustment module is used to match the processed video image with the processed multiple pictures and sounds by using a data matching algorithm, and synthesize the pictures and audios for which matching results exist; for pictures and sounds for which matching results do not exist, perform audio adjustment on the sounds based on an audio waveform adjustment algorithm, and synthesize the adjusted audio with the video image; The video detection module detects the lens image and audio playback effect simultaneously based on the video detection algorithm; The early warning broadcast module is used to detect the video images that have passed the detection by the video detection module based on the early warning broadcast algorithm, and to issue an early warning broadcast when the patient has abnormal behavior or abnormal state.