A gynecological and obstetrical baby care system and method
By placing infrared cameras in the neonatal care room, collecting facial image frame sequences and performing micro-expression key frame extraction and spatiotemporal feature analysis, and combining them with sound signals, intelligent and precise obstetric and gynecological infant care is achieved, solving the problems of subjectivity in infant care and insufficient comprehensive analysis of multimodal information in existing technologies.
Patent Information
- Application Number
- CN202510288289.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing obstetrics and gynecology infant care mainly relies on manual monitoring, which has problems of strong subjectivity and low accuracy. In addition, existing computer vision technology cannot effectively adapt to the particularity of infants' facial expressions and cannot comprehensively analyze multimodal information.
An infrared camera is placed in the neonatal care room to collect facial image frame sequences. Through micro-expression key frame extraction and spatiotemporal feature extraction, facial image collection is performed using the infrared camera, combined with sound signal analysis to achieve intelligent care response.
It improves the accuracy and intelligence of infant care, can more accurately meet the actual needs of infants and reduce misjudgments.
Smart Images

Figure CN120220995B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of safety monitoring, and in particular to a gynecological and obstetrics infant care system and method. Background Art
[0002] Existing infant care in obstetrics and gynecology departments primarily relies on manual monitoring by nurses or family members. This approach typically assesses the infant's condition by observing their facial expressions, crying, and movements, and then makes informed decisions based on experience. However, this approach is highly subjective and easily influenced by the caregiver's experience, attention, and environmental factors, resulting in unstable judgments.
[0003] Existing research has attempted to use computer vision technology to recognize infant facial expressions. However, these methods are typically based on adult facial expression recognition models and cannot effectively adapt to the specific characteristics of infant facial expressions, such as the narrow range of expression variation and the limited number of expression categories. Furthermore, existing emotion recognition systems often rely on data from a single modality, such as facial images or crying sounds, and fail to comprehensively analyze the infant's multimodal information. This results in limited recognition accuracy and an inability to provide precise care response plans. Summary of the Invention
[0004] The present application provides a gynecological and obstetric infant care system and method, which is used to solve the technical problems in the existing technology that multimodal care information cannot be integrated efficiently and accurately, and the care is not well aligned with the actual situation.
[0005] In view of the above problems, the present application provides a system and method for caring for infants in obstetrics and gynecology.
[0006] In a first aspect of the present application, a gynecological and obstetric baby care system is provided, the system comprising:
[0007] An infrared camera is placed in the neonatal care room, and facial images are collected using the infrared camera to obtain a facial image frame sequence;
[0008] Traversing the facial image frame sequence to extract micro-expression key frames to obtain a micro-expression key frame sequence;
[0009] Extracting spatiotemporal features from the micro-expression key frame sequence to obtain a micro-expression feature vector sequence;
[0010] A nursing response plan is matched according to the micro-expression feature vector sequence to obtain a target nursing response plan.
[0011] The second aspect of the present application provides a method for caring for infants in obstetrics and gynecology, the method comprising:
[0012] A facial image frame sequence acquisition module is used to deploy an infrared camera in the neonatal care room, use the infrared camera to collect facial images, and obtain a facial image frame sequence;
[0013] A micro-expression key frame sequence acquisition module is used to traverse the facial image frame sequence to extract micro-expression key frames and obtain a micro-expression key frame sequence;
[0014] A micro-expression feature vector sequence acquisition module is used to extract spatiotemporal features of the micro-expression key frame sequence to obtain a micro-expression feature vector sequence;
[0015] The target nursing response solution acquisition module is used to match the nursing response solution according to the micro-expression feature vector sequence to obtain the target nursing response solution.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0017] This application deploys an infrared camera in a neonatal care room, uses the infrared camera to capture facial images, obtains a facial image frame sequence, then traverses the facial image frame sequence to extract micro-expression key frames, obtains a micro-expression key frame sequence, further extracts spatiotemporal features from the micro-expression key frame sequence, obtains a micro-expression feature vector sequence, and then matches a care response plan based on the micro-expression feature vector sequence to obtain a target care response plan. This achieves the technical effect of improving infant care quality and obtaining a care plan that meets actual care needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Attachment Figure 1 This is a schematic structural diagram of a gynecology and obstetrics baby care system provided by an embodiment of the present invention.
[0019] Attachment Figure 2 The present invention provides a flow chart of a method for caring for infants in obstetrics and gynecology.
[0020] Reference numerals shown in the accompanying drawings:
[0021] Facial image frame sequence acquisition module 11, micro-expression key frame sequence acquisition module 12, micro-expression feature vector sequence acquisition module 13, target care response plan acquisition module 14. DETAILED DESCRIPTION
[0022] The application will be further described below in connection with specific embodiments. It should be understood that these embodiments are only used to explain the application and not intended to limit the scope of the application. Furthermore, it should be understood that after reading the content of the application, those skilled in the art can make various modifications or changes to the application, and these equivalent forms also fall within the scope defined by the appended claims. It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] Embodiment one, as shown in the accompanying drawings, the present application provides a gynecological and obstetric baby care system, the system comprises: Figure 1
[0024] The face image frame sequence obtaining module 11 is configured to arrange an infrared camera in the neonatal care room, and collect face images by using the infrared camera to obtain a face image frame sequence.
[0025] In one possible embodiment, an infrared camera is arranged in the neonatal care room. Preferably, a high-resolution and low-noise infrared camera is selected to ensure that the baby's facial expressions can be clearly captured in a low-light environment. When arranging, the camera should avoid direct sunlight as much as possible, and be installed above or beside the baby's bed to obtain stable and complete face images. The infrared camera uses a real-time video stream acquisition method to continuously acquire face image frame sequences at a fixed frame rate (such as 30 fps). Considering the head movement of the baby, to ensure data quality,
[0026] For example, assuming that the system acquires video at a frame rate of 30 fps, 30 face images are extracted per second to obtain the face image frame sequence. Preferably, the system can set a time window (such as 5 seconds) and cache 150 images for subsequent micro-expression analysis. The face image frame sequence serves as the basis for subsequent micro-expression key frame extraction and emotion recognition, ensuring data quality while providing stable input for subsequent analysis.
[0027] The micro-expression key frame sequence obtaining module 12 is configured to traverse the face image frame sequence to extract micro-expression key frames and obtain a micro-expression key frame sequence.
[0028] Further, the micro-expression key frame sequence obtaining module 12 is configured to perform the following steps:
[0029] traversing the face image frame sequence to perform grayscale processing to obtain a face image grayscale frame sequence;
[0030] Calculating the amplitudes of optical flow field vectors of adjacent frames of the facial image grayscale frame sequence based on the Farneback optical flow method to obtain an optical flow field vector amplitude change curve;
[0031] Performing a nearest neighbor clustering analysis on the optical flow field vector amplitude change curve to obtain an optical flow field vector amplitude cluster;
[0032] Micro-expression key frames are extracted from the facial image grayscale frame sequence based on the optical flow field vector amplitude cluster to obtain the micro-expression key frame sequence.
[0033] In one possible embodiment, since the facial image frame sequence is a continuous image and contains a large number of redundant images, in order to accurately analyze the changes in the baby's facial micro-expressions reflected in the image, it is necessary to extract micro-expression key frames from the facial image frame sequence to obtain the micro-expression key frame sequence.
[0034] First, the facial image frame sequence is traversed and grayscaled to remove color information and reduce computational complexity, resulting in a sequence of facial image grayscale frames. Then, the optical flow field between adjacent frames is calculated using the Farneback optical flow method. The motion vectors of each pixel are extracted and their amplitudes are calculated, resulting in an optical flow vector amplitude change curve. Next, this curve is processed using nearest neighbor clustering analysis to identify regions of significant variation and form optical flow vector amplitude clusters. These clusters reflect regions of subtle changes in facial expression. Finally, based on these amplitude clusters, key frames are extracted from the facial image grayscale frame sequence to form a micro-expression key frame sequence for subsequent infant emotion feature analysis and care plan matching. This process aims to accurately extract key moments of infant micro-expressions, reduce redundant frames, and improve the accuracy and computational efficiency of subsequent emotion recognition and response plan matching.
[0035] In one possible embodiment, the optical flow field vector amplitude change curve is a temporal trend curve of the optical flow field vector amplitude, and is used to detect grayscale frames of facial images in which micro-expressions occur. A facial image grayscale frame sequence refers to a sequence of continuous image frames that have undergone grayscale processing, wherein each frame is a grayscale image, retaining only the brightness information of facial features and removing color interference.
[0036] Preferably, each frame of the collected facial image frames is usually composed of three channels: red (R), green (G), and blue (B), and the core of grayscale is to merge these three channels in some way while ensuring image clarity. For example, the weighted average method (Gray = 0.299R + 0.587G + 0.114B) is used to convert color information to obtain a single-channel grayscale image. By performing grayscale processing, the facial image grayscale frame sequence is obtained. This step can reduce data complexity and reduce the amount of calculation, while enhancing the recognizability of key information such as facial contours and detailed textures, providing more stable input data for subsequent optical flow calculations. In addition, grayscale can eliminate the interference caused by color changes, making subsequent micro-expression recognition more accurate.
[0037] Furthermore, the micro-expression key frame sequence acquisition module 12 is configured to perform the following steps:
[0038] Based on the Farneback optical flow method, dense optical flow calculation is performed on adjacent frames in the facial image grayscale frame sequence in a time-ordered order to obtain an adjacent frame optical flow field sequence, wherein each adjacent frame optical flow field in the adjacent frame optical flow field sequence is an optical flow field between a current frame and a previous frame;
[0039] Traversing the adjacent frame optical flow field sequence to calculate the horizontal component and the vertical component, and obtaining the adjacent frame optical flow field horizontal component set sequence and the adjacent frame optical flow field vertical component set sequence;
[0040] Using the optical flow field vector amplitude calculation function, the horizontal component set sequence of the adjacent frame optical flow field and the vertical component set sequence of the adjacent frame optical flow field are calculated to obtain the optical flow field vector amplitude sequence;
[0041] The optical flow field vector amplitude sequence is fitted to construct the optical flow field vector amplitude change curve.
[0042] Furthermore, the optical flow field vector amplitude calculation function is:
[0043]
[0044] Among them, M ROI is the magnitude of the optical flow field vector, N is the number of pixels in the grayscale frame of the facial image, N is a positive integer, dx(x i ,y i ) is the horizontal component of the adjacent frame optical flow field of the i-th pixel in the grayscale frame of the facial image, dy(x i ,y i ) is the vertical component of the adjacent frame optical flow field of the i-th pixel in the grayscale frame of the facial image.
[0045] In one possible embodiment, the Farneback optical flow method is a dense optical flow estimation algorithm that can calculate pixel-level motion information between adjacent frames in a video sequence. The Farneback optical flow method can provide an optical flow vector for each pixel and is suitable for global motion analysis. The adjacent frame optical flow field vector amplitude calculation is to calculate the motion vector of each pixel between two adjacent frames of images. The vector consists of a horizontal component (dx) and a vertical component (dy). The adjacent frame optical flow field sequence refers to the motion vector field of all pixels in the current frame and the previous frame, where the motion of each pixel is represented by a vector (dx, dy), and the magnitude and direction of the vector represent the motion of the pixel. Among them, each adjacent frame optical flow field in the adjacent frame optical flow field sequence is the optical flow field between the current frame and the previous frame. The optical flow field vector amplitude change curve represents the trend of the motion intensity of the facial area over time in the entire image frame sequence, and can be used to detect subtle facial movements, such as the occurrence of micro-expressions.
[0046] The horizontal component (dx) and vertical component (dy) of the optical flow field between adjacent frames represent the horizontal and vertical motion of a pixel, respectively. dx is the x-component of motion, and dy is the y-component of motion. The optical flow vector magnitude measures the overall motion intensity of a grayscale frame of a facial image. It is calculated by averaging the optical flow vector magnitudes of all pixels.
[0047] Preferably, any two adjacent frames of the facial image grayscale frame sequence are extracted, and each two frames are used as a group of inputs, and Farneback optical flow calculation is performed in chronological order. Specifically, for the current frame and the previous frame, the motion vectors (dx, dy) of all pixels in the image are calculated to generate the optical flow field of the current frame relative to the previous frame. This process continues until all adjacent frames are calculated, thereby obtaining a complete sequence of adjacent frame optical flow fields. By performing running vector analysis, the motion information of the facial area, especially small muscle movements, is extracted. For example, when a baby smiles, the muscles at the corners of the mouth and cheeks will rise slightly, and this change can be captured by the vector direction and size in the optical flow field. By analyzing the entire optical flow field sequence, the movement pattern of the facial muscles can be identified, providing basic data for the subsequent extraction of micro-expression key frames.
[0048] The system sequentially traverses each frame in the optical flow sequence and decomposes the motion vectors of all pixels therein, extracting the dx component in the x direction and the dy component in the y direction. The dx and dy components of each frame form a set. Over time, these sets are arranged in sequence to form a complete sequence of sets of horizontal components of the optical flow field in adjacent frames and a complete sequence of sets of vertical components of the optical flow field in adjacent frames.
[0049] By breaking down complex optical flow information into independent horizontal and vertical motion features, subsequent motion pattern analysis is facilitated. For example, when an infant frowns due to discomfort, this primarily occurs in the vertical direction of the forehead (with a negative dy value); whereas when a baby displays a happy expression, the corners of their mouth may stretch laterally, primarily in the horizontal direction (with a positive dx value). This decomposition helps more accurately capture and analyze subtle changes in facial micro-expressions, providing data support for subsequent micro-expression keyframe extraction.
[0050] Preferably, the optical flow vector amplitude calculation function is used to quantify the intensity of motion in the facial region. By calculating the motion vector amplitudes for all pixels in each frame, a time-varying sequence of optical flow vector amplitudes is obtained. The optical flow vector amplitude variation curve is a smooth curve obtained by fitting the optical flow vector amplitude sequence. It is used to describe the temporal variation trend of the optical flow vector amplitude and reflect the dynamic characteristics of facial expression movement. Optionally, the optical flow vector amplitude sequence is fitted using a moving average method, with local average calculation performed using a sliding window pre-set by a person skilled in the art to smooth short-term fluctuations and make the curve more stable.
[0051] For example, when the optical flow vector amplitude sequence includes 10 optical flow vector amplitudes, namely: 0.2, 0.3, 0.5, 0.8, 1.2, 1.5, 1.0, 0.6, 0.3, 0.1. A person skilled in the art pre-sets the sliding window to 3. At this time, performing the sliding window calculation, the smoothed data point sequence obtained is: 0.33, 0.53, 0.83, 1.17, 1.23, 1.03, 0.63, 0.33. Then, a coordinate system is constructed with time as the horizontal axis and the optical flow vector amplitude as the vertical axis. The smoothed data point sequence is sequentially input into the coordinate system to obtain a coordinate point sequence, and the coordinate point sequence is sequentially connected to obtain the fitted optical flow vector amplitude change curve.
[0052] Furthermore, the micro-expression key frame sequence acquisition module 12 is configured to perform the following steps:
[0053] Extracting the local maximum value in the optical flow field vector amplitude change curve to obtain a set of neighbor cluster center points;
[0054] Taking the center point sets of the nearest clusters as starting points, the nearest cluster analysis is performed according to a preset amplitude difference threshold to obtain the optical flow field vector amplitude clusters.
[0055] In one embodiment of the present application, a local maximum refers to a point in a certain neighborhood whose value is greater than the values of other points in the neighborhood, indicating a significant peak point in the optical flow field vector amplitude change curve. The neighboring cluster center point is a reference for clustering. Combined with a preset amplitude difference threshold, the neighborhood belonging to each neighboring cluster center point can be delineated, thereby obtaining a set of optical flow field vectors for each neighboring cluster center point, and summarizing them to obtain the optical flow field vector amplitude cluster. Trend-based filtering refers to filtering out amplitude clusters that conform to trend changes by calculating the mean of the amplitude clusters, and excluding outliers or noise points. The preset amplitude threshold is the minimum amplitude difference between the optical flow field vector amplitude and the neighboring cluster center point when it can be demarcated into the neighborhood of the neighboring cluster center point, which is pre-set by those skilled in the art. It is used to screen valid micro-expression key frames to ensure that the selected key frames have sufficient motion amplitude and represent micro-expression changes.
[0056] Preferably, the optical flow field vector amplitude change curve is traversed to extract local maxima, and the coordinate points where the extracted local maxima are located are used as the center points of the nearest clusters, thereby obtaining the set of nearest cluster center points. The set of nearest cluster center points is composed of coordinate points whose amplitudes are greater than the amplitudes of the optical flow field vectors of the two adjacent left and right coordinate points. Then, taking each nearest cluster center point as the starting point, the amplitude difference between it and the adjacent points is calculated. If the difference is less than the preset amplitude difference threshold, the point is included in the same cluster until a point is encountered that drops significantly or exceeds the threshold range, thereby forming the optical flow field vector amplitude cluster.
[0057] Furthermore, the micro-expression key frame sequence acquisition module 12 is configured to perform the following steps:
[0058] Traversing and calculating the mean of the optical flow field vector amplitude clusters to obtain an optical flow field vector amplitude mean set;
[0059] Taking the optical flow field vector amplitude mean value set as the trend starting point, performing central trend screening on the optical flow field vector amplitude clusters to obtain screened optical flow field vector amplitude clusters;
[0060] Performing mean calculation on the filtered optical flow field vector amplitude clusters to obtain a filtered optical flow field vector amplitude mean set;
[0061] Determine whether the filtered optical flow field vector amplitude mean set is greater than or equal to a preset amplitude threshold; if so, sort the multiple facial image grayscale frames corresponding to the optical flow field vector amplitude cluster in chronological order to obtain the micro-expression key frame sequence.
[0062] In a possible embodiment, the mean values of the optical flow field vector magnitude clusters are calculated to obtain a set of optical flow field vector magnitude mean values. The set of optical flow field vector magnitude mean values reflects the average level of the set of optical flow field vector magnitudes of each near neighbor cluster center in the optical flow field vector magnitude cluster. Then, the set of optical flow field vector magnitude mean values is taken as the trend starting point, and a trend starting point neighborhood is constructed according to a preset screening step size to obtain a set of trend starting point neighborhoods. Each trend starting point neighborhood corresponds to an optical flow field vector magnitude mean value. The preset screening step size is the magnitude of a single expansion for centralized trend screening preset by a person skilled in the art.
[0063] Preferably, the number of optical flow field vector magnitudes in the set of trend starting point neighborhoods is counted respectively, and the counting result is compared with twice the preset screening step size to obtain the density of the set of trend starting point neighborhoods, that is, a set of trend neighborhood densities. Then, the left and right ends of the set of trend starting point neighborhoods are diffused outward according to the preset screening step size to obtain a set of trend starting point first diffusion neighborhoods. Based on the same principle as obtaining the set of trend neighborhood densities, the density of the set of trend starting point first diffusion neighborhoods is calculated to obtain a set of trend starting point first diffusion neighborhood densities.
[0064] It is determined whether the set of trend starting point first diffusion neighborhood densities is greater than or equal to the corresponding set of trend neighborhood densities. If yes, the set of trend starting point first diffusion neighborhoods is diffused outward again based on the preset screening step size to obtain a set of trend starting point second diffusion neighborhoods, and the diffusion is continued until the neighborhood density difference between two adjacent diffusion is less than or equal to a preset difference value, the diffusion is stopped, and a set of trend starting point target diffusion neighborhoods is obtained. The set of trend starting point target diffusion neighborhoods is summarized to obtain the screened optical flow field vector magnitude cluster.
[0065] Then, the mean values of the screened optical flow field vector magnitude clusters are calculated to obtain a set of screened optical flow field vector magnitude mean values that can reflect the magnitude of each screened optical flow field vector magnitude set. The preset magnitude threshold is the minimum magnitude corresponding to a micro-expression key frame preset by a person skilled in the art. It is determined whether the set of screened optical flow field vector magnitude mean values is greater than or equal to the preset magnitude threshold. If yes, it indicates that the corresponding facial image grayscale frame is a micro-expression key frame. At this time, the multiple facial image grayscale frames corresponding to the optical flow field vector magnitude cluster are sorted in chronological order to obtain the sequence of micro-expression key frames.
[0066] The micro-expression feature vector sequence obtaining module 13 is configured to perform spatiotemporal feature extraction on the sequence of micro-expression key frames to obtain a sequence of micro-expression feature vectors.
[0067] In one possible embodiment, a spatiotemporal feature extractor is constructed, and the spatiotemporal feature extractor is used to perform feature extraction on the micro-expression key frame sequence to obtain a micro-expression feature vector sequence. The spatiotemporal feature extractor refers to a model or algorithm that can simultaneously capture temporal (Temporal) and spatial (Spatial) information. In micro-expression analysis, spatial features refer to the local texture and shape changes of the face in a single frame image (such as the slight upward movement of the corners of the mouth and the changes in fine wrinkles around the eyes), while temporal features refer to the changes in facial dynamics between multiple frames (such as the start, development, and end of micro-expressions). The micro-expression feature vector sequence is a numerical high-dimensional vector representation converted from the key frame sequence after feature extraction, which can be used for subsequent classification, analysis and other tasks.
[0068] Preferably, a plurality of sample micro-expression key frames and a plurality of sample micro-expression feature vectors are obtained as training data, and the framework constructed based on the feedforward neural network is supervisedly trained using the training data, and the network parameters of the framework are updated according to the output results during training until the training converges, thereby obtaining the trained spatiotemporal feature extractor.
[0069] The target nursing response solution acquisition module 14 is used to match nursing response solutions according to the micro-expression feature vector sequence to obtain a target nursing response solution.
[0070] Furthermore, the target nursing response solution obtaining module 14 is configured to perform the following steps:
[0071] Using a microphone placed at the target infant's nursing bed to collect sound signals, a sound signal sequence is obtained;
[0072] Extracting abnormal features from the sound signal sequence to obtain a set of abnormal sound features;
[0073] The target nursing response plan is modified based on the abnormal sound feature set to obtain a modified nursing response plan.
[0074] In one embodiment, a targeted care response plan refers to an intervention measure taken in response to a specific state of a care recipient (e.g., an infant), such as adjusting ambient lighting, playing soothing music, or notifying a caregiver. The micro-expression feature vector sequence generated by the previous step represents the dynamic features of the infant's facial micro-expressions. These features can be used to infer the infant's emotions or needs (e.g., restlessness, hunger, sleepiness).
[0075] Preferably, a pre-built support vector machine (SVM) is used to classify micro-expression features, for example, to distinguish states such as "mild anxiety", "anxious crying", and "comfortable relaxation". Then, the plan that best matches the current micro-expression state is retrieved from the predefined care plan database. For example, the care plan corresponding to mild anxiety is slight cradle vibration + soft background music, the care plan corresponding to anxious crying is to increase the cradle vibration amplitude + simulate the mother's heartbeat, and the care plan corresponding to sleepiness is to turn off the strong light source + provide a warm wrapping feeling. The care plan database is obtained by those skilled in the art setting the care plans corresponding to different micro-expression feature vectors according to actual conditions and storing them in the database. Then, the matched target care response plan is used as output to guide intelligent care equipment (such as smart cribs, automatic soothing systems) to take corresponding measures. The role of this process is to realize automated and intelligent baby status monitoring and soothing based on facial micro-expressions, thereby improving the accuracy and response efficiency of care.
[0076] In one embodiment, sound signal collection refers to recording the sounds around the target baby, such as crying, laughing, babbling or breathing, through a microphone. The sound signal sequence is to collect these sounds and store them in the form of time series data for subsequent analysis. Abnormal feature extraction mainly refers to feature analysis of sound signals to identify abnormal situations, such as the frequency, intensity, duration, etc. of crying. The sound abnormality feature set is an abnormal pattern extracted from the collected sound signal, such as a sudden increase in crying, intermittent sobbing, a long period of silence, etc., which may indicate that the baby has needs or discomfort. Modifying the care response plan means adjusting the plan based on the abnormal sound features on the basis of the existing care plan to more accurately meet the needs of the baby, such as increasing the intensity of comfort or notifying the caregiver.
[0077] Preferably, a sound abnormality feature extractor is constructed using a convolutional neural network. By obtaining multiple sample sound signal sequences and corresponding multiple sample sound abnormality feature sets as a sample training data set, the sample training data set is equally divided into n groups of sample training data. Then, the framework constructed based on the convolutional neural network is supervised and trained using the n groups of sample training data in sequence. The network parameters of the framework are modified according to the output results of each group of sample training data until the training converges, thereby obtaining the trained sound abnormality feature extractor. The trained sound abnormality feature extractor is used to extract abnormal features from the sound signal sequence to obtain the sound abnormality feature set.
[0078] Preferably, an existing baby care database is searched and the optimal care measures corresponding to the current abnormal sound feature set are analyzed. For example: high-frequency rapid crying + long-term persistence may indicate discomfort (such as pain, eczema). Based on the matching results, the target care response plan is modified, such as: increasing the intensity of the soothing action (such as increasing the swing amplitude of the cradle), changing the soothing mode (such as switching from light music to simulating the mother's heartbeat), etc., to obtain the modified care response plan. The dynamic adaptation of the care plan is achieved, so that it can be adjusted according to the baby's immediate state, improving the accuracy of intelligent care, and reducing the technical effect of misjudgment.
[0079] In summary, the embodiments of the present application have at least the following technical effects:
[0080] This application deploys an infrared camera in a neonatal care room, uses the infrared camera to capture facial images, obtains a facial image frame sequence, then traverses the facial image frame sequence to extract micro-expression key frames, obtains a micro-expression key frame sequence, further extracts spatiotemporal features from the micro-expression key frame sequence, obtains a micro-expression feature vector sequence, and then matches a care response plan based on the micro-expression feature vector sequence to obtain a target care response plan. This achieves the technical effect of improving infant care quality and obtaining a care plan that meets actual care needs.
[0081] Embodiment 2 is based on the same inventive concept as the obstetrics and gynecology baby care system in the above embodiment, as shown in the attached Figure 2 As shown, the present application provides a method for caring for infants in obstetrics and gynecology. The method and system embodiments in the present application are based on the same inventive concept. The method includes:
[0082] An infrared camera is placed in the neonatal care room, and facial images are collected using the infrared camera to obtain a facial image frame sequence;
[0083] Traversing the facial image frame sequence to extract micro-expression key frames to obtain a micro-expression key frame sequence;
[0084] Extracting spatiotemporal features from the micro-expression key frame sequence to obtain a micro-expression feature vector sequence;
[0085] A nursing response plan is matched according to the micro-expression feature vector sequence to obtain a target nursing response plan.
[0086] Furthermore, the method further comprises:
[0087] Traversing the facial image frame sequence and performing grayscale processing to obtain a facial image grayscale frame sequence;
[0088] Calculating the amplitudes of optical flow field vectors of adjacent frames of the facial image grayscale frame sequence based on the Farneback optical flow method to obtain an optical flow field vector amplitude change curve;
[0089] Performing a nearest neighbor clustering analysis on the optical flow field vector amplitude change curve to obtain an optical flow field vector amplitude cluster;
[0090] Micro-expression key frames are extracted from the facial image grayscale frame sequence based on the optical flow field vector amplitude cluster to obtain the micro-expression key frame sequence.
[0091] Furthermore, the method further comprises:
[0092] Based on the Farneback optical flow method, dense optical flow calculation is performed on adjacent frames in the facial image grayscale frame sequence in a time-ordered order to obtain an adjacent frame optical flow field sequence, wherein each adjacent frame optical flow field in the adjacent frame optical flow field sequence is an optical flow field between a current frame and a previous frame;
[0093] Traversing the adjacent frame optical flow field sequence to calculate the horizontal component and the vertical component, and obtaining the adjacent frame optical flow field horizontal component set sequence and the adjacent frame optical flow field vertical component set sequence;
[0094] Using the optical flow field vector amplitude calculation function, the horizontal component set sequence of the adjacent frame optical flow field and the vertical component set sequence of the adjacent frame optical flow field are calculated to obtain the optical flow field vector amplitude sequence;
[0095] The optical flow field vector amplitude sequence is fitted to construct the optical flow field vector amplitude change curve.
[0096] Furthermore, the optical flow field vector amplitude calculation function is:
[0097]
[0098] Among them, M ROI is the magnitude of the optical flow field vector, N is the number of pixels in the grayscale frame of the facial image, N is a positive integer, dx(x i ,y i ) is the horizontal component of the adjacent frame optical flow field of the i-th pixel in the grayscale frame of the facial image, dy(x i ,y i ) is the vertical component of the adjacent frame optical flow field of the i-th pixel in the grayscale frame of the facial image.
[0099] Furthermore, the method further comprises:
[0100] Extracting the local maximum value in the optical flow field vector amplitude change curve to obtain a set of neighbor cluster center points;
[0101] Taking the center point sets of the nearest clusters as starting points, the nearest cluster analysis is performed according to a preset amplitude difference threshold to obtain the optical flow field vector amplitude clusters.
[0102] Furthermore, the method further comprises:
[0103] Traversing and calculating the mean of the optical flow field vector amplitude clusters to obtain an optical flow field vector amplitude mean set;
[0104] Taking the optical flow field vector amplitude mean value set as the trend starting point, performing central trend screening on the optical flow field vector amplitude clusters to obtain screened optical flow field vector amplitude clusters;
[0105] Performing mean calculation on the filtered optical flow field vector amplitude clusters to obtain a filtered optical flow field vector amplitude mean set;
[0106] Determine whether the filtered optical flow field vector amplitude mean set is greater than or equal to a preset amplitude threshold; if so, sort the multiple facial image grayscale frames corresponding to the optical flow field vector amplitude cluster in chronological order to obtain the micro-expression key frame sequence.
[0107] Furthermore, the method further comprises:
[0108] Using a microphone placed at the target infant's nursing bed to collect sound signals, a sound signal sequence is obtained;
[0109] Extracting abnormal features from the sound signal sequence to obtain a set of abnormal sound features;
[0110] The target nursing response plan is modified based on the abnormal sound feature set to obtain a modified nursing response plan.
[0111] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0112] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
[0113] The specification and drawings are, of course, to be regarded in an illustrative rather than a restrictive sense. It is to be understood that any such modifications, variations, combinations or equivalents that fall within the scope of the application are intended to be embraced herein.
Claims
1. A gynecological and obstetric baby care system, characterized in that: The system comprises: A facial image frame sequence acquisition module is used to deploy an infrared camera in the neonatal care room, use the infrared camera to collect facial images, and obtain a facial image frame sequence; A micro-expression key frame sequence acquisition module is used to traverse the facial image frame sequence to extract micro-expression key frames and obtain a micro-expression key frame sequence; A micro-expression feature vector sequence acquisition module is used to extract spatiotemporal features of the micro-expression key frame sequence to obtain a micro-expression feature vector sequence; a target nursing response solution acquisition module, configured to match nursing response solutions according to the micro-expression feature vector sequence to obtain a target nursing response solution; The micro-expression key frame sequence acquisition module is used to perform the following steps: Traversing the facial image frame sequence and performing grayscale processing to obtain a facial image grayscale frame sequence; Calculating the amplitude of the optical flow field vector of adjacent frames of the facial image grayscale frame sequence based on the Farneback optical flow method to obtain an optical flow field vector amplitude change curve; Performing a nearest neighbor clustering analysis on the optical flow field vector amplitude change curve to obtain an optical flow field vector amplitude cluster; Extracting micro-expression key frames from the facial image grayscale frame sequence based on the optical flow field vector amplitude cluster to obtain the micro-expression key frame sequence; The method further includes performing a nearest neighbor clustering analysis on the optical flow field vector amplitude change curve to obtain an optical flow field vector amplitude cluster. Extracting the local maximum value in the optical flow field vector amplitude change curve to obtain a set of neighboring cluster center points, wherein the set of neighboring cluster center points are coordinate points whose amplitudes are greater than the amplitudes of the optical flow field vectors of two adjacent left and right coordinate points; Taking the set of neighbor cluster center points as the starting point, neighbor cluster analysis is performed according to the preset amplitude difference threshold to obtain the optical flow field vector amplitude cluster, wherein the neighbor cluster analysis specifically takes each neighbor cluster center point as the starting point, calculates the amplitude difference between it and the neighboring points, and if the difference is less than the preset amplitude difference threshold, then the point is included in the same cluster until a point is encountered that drops significantly or exceeds the threshold range; The method further comprises: extracting micro-expression key frames from the facial image grayscale frame sequence based on the optical flow field vector amplitude cluster to obtain the micro-expression key frame sequence. Traversing and calculating the mean of the optical flow field vector amplitude clusters to obtain an optical flow field vector amplitude mean set; Taking the optical flow field vector amplitude mean value set as the trend starting point respectively, constructing the trend starting point neighborhood according to the preset screening step size to obtain the trend starting point neighborhood set; Counting the number of optical flow field vector amplitudes in the trend starting point neighborhood set respectively, and multiplying the statistical result by twice the preset screening step length to obtain a trend neighborhood density set; Diffusion of the left and right ends of the trend starting point neighborhood set outward according to a preset screening step length to obtain a trend starting point first diffusion neighborhood set and a trend starting point first diffusion neighborhood density set; Determine whether the primary diffusion neighborhood density set of the trend starting point is greater than or equal to the corresponding trend neighborhood density set. If so, diffuse the primary diffusion neighborhood set of the trend starting point outward again based on the preset screening step size to obtain the secondary diffusion neighborhood set of the trend starting point, and continue to diffuse until the difference in neighborhood density between two adjacent diffusions is less than or equal to the preset difference, stop the diffusion, obtain the target diffusion neighborhood set of the trend starting point, and summarize the target diffusion neighborhood set of the trend starting point to obtain the filtered optical flow field vector amplitude cluster; Performing mean calculation on the filtered optical flow field vector amplitude clusters to obtain a filtered optical flow field vector amplitude mean set; Determine whether the filtered optical flow field vector amplitude mean set is greater than or equal to a preset amplitude threshold; if so, sort the multiple facial image grayscale frames corresponding to the optical flow field vector amplitude cluster in chronological order to obtain the micro-expression key frame sequence.
2. The obstetrics and gynecology baby care system according to claim 1, characterized in that: The micro-expression key frame sequence acquisition module is used to perform the following steps: Based on the Farneback optical flow method, dense optical flow calculation is performed on adjacent frames in the facial image grayscale frame sequence in a time-ordered order to obtain an adjacent frame optical flow field sequence, wherein each adjacent frame optical flow field in the adjacent frame optical flow field sequence is an optical flow field between a current frame and a previous frame; Traversing the adjacent frame optical flow field sequence to calculate the horizontal component and the vertical component, and obtaining the adjacent frame optical flow field horizontal component set sequence and the adjacent frame optical flow field vertical component set sequence; Using the optical flow field vector amplitude calculation function, the horizontal component set sequence of the adjacent frame optical flow field and the vertical component set sequence of the adjacent frame optical flow field are calculated to obtain the optical flow field vector amplitude sequence; The optical flow field vector amplitude sequence is fitted to construct the optical flow field vector amplitude change curve.
3. The obstetrics and gynecology baby care system according to claim 2, characterized in that: The optical flow field vector amplitude calculation function is: ; in, is the optical flow field vector amplitude, is the number of pixels in the grayscale frame of the facial image, N is a positive integer, is the grayscale frame of the facial image The horizontal component of the optical flow field of adjacent frames of pixels, is the grayscale frame of the facial image The vertical component of the optical flow field of adjacent frames of pixels.
4. The obstetrics and gynecology baby care system according to claim 1, characterized in that: The target care response scheme acquisition module is used to perform the following steps: Using a microphone placed at the target infant's nursing bed to collect sound signals, a sound signal sequence is obtained; Extracting abnormal features from the sound signal sequence to obtain a set of abnormal sound features; The target nursing response plan is modified based on the abnormal sound feature set to obtain a modified nursing response plan.
5. A method for caring for infants in obstetrics and gynecology, characterized in that: The method is performed by an obstetrics and gynecology baby care system according to any one of claims 1 to 4, and the method comprises: An infrared camera is placed in the neonatal care room, and facial images are collected using the infrared camera to obtain a facial image frame sequence; Traversing the facial image frame sequence to extract micro-expression key frames to obtain a micro-expression key frame sequence; Extracting spatiotemporal features from the micro-expression key frame sequence to obtain a micro-expression feature vector sequence; A nursing response plan is matched according to the micro-expression feature vector sequence to obtain a target nursing response plan.
Citation Information
Patent Citations
Infant monitoring method and device, computer equipment and storage medium
CN109800646A