Control system and controlled device

By integrating urine sensing, sound recognition, and video monitoring into the control system, the problems of traditional baby cradles being limited in movement and lacking in intelligence have been solved. This has enabled multi-mode control and adaptive care functions for intelligent baby cradles, thus improving the level of intelligence in baby care.

CN120544598BActive Publication Date: 2026-04-03YUYAO DANENG MOULD TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional baby cradles have a single movement pattern and cannot simulate the complex movement characteristics of natural cuddling. They also have low levels of intelligence, lack interactive and behavioral recognition capabilities, and cannot meet the comprehensive care and real-time monitoring needs of modern families.

Method used

The system employs a control system, including a data acquisition module, a controlled module, and an execution module. Through urine sensing, sound recognition, video monitoring, and touch/remote/voice control, combined with Mel-frequency cepstral coefficient analysis, deep hidden Markov models, and context-constrained word segmentation models, it can detect infant crying and identify sleep states, generating corresponding control commands to drive the movement of the cradle and play audio.

Benefits of technology

It enhances the intelligence of baby cradles, featuring integrated control via touch, remote, voice, and automation. It can automatically detect bedwetting and crying, and adaptively adjust cradle movement and audio playback, thereby improving the intelligence and real-time monitoring capabilities of baby care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544598B_ABST
    Figure CN120544598B_ABST
Patent Text Reader

Abstract

This invention discloses a control system and a controlled device, relating to the field of intelligent control, including a data acquisition module, a controlled module, an automatic control module, and an execution module. The data acquisition module senses urine contact and generates a high-level signal, acquiring audio digital signal sequences and video digital signal sequences. The controlled module senses touch or receives remote control signals and generates control commands, generating prompts based on the audio digital signal sequences for infant crying or human language selection, or analyzing voice content through speech recognition technology to generate control commands. The automatic control module periodically captures video digital signal sequences, identifies the infant's sleep state through transcoding, facial recognition, and optical flow detection, and automatically generates control commands or prompts, continuously generating prompts during high-level signal reception. The execution module executes the control commands, receives the video digital signal sequences and prompts, and uploads them to a cloud platform, realizing remote and intelligent collaborative control of the cradle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control, and more particularly to a control system and a controlled device. Background Technology

[0002] Traditional baby cradles have limited functionality, only capable of simple rocking motions, and cannot meet the needs of modern families for comprehensive baby care and real-time monitoring.

[0003] Existing electric cradles have the following drawbacks: they have a single movement mode and cannot simulate the complex movement characteristics of a natural embrace; they have a low level of intelligence, relying on preset programs or close-range human intervention, and lack interactive and behavioral recognition capabilities. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies by proposing a control system and a controlled device to realize an intelligent baby cradle with four control modes: button, remote, voice, and automatic.

[0005] The technical solution to achieve the purpose of this invention is as follows:

[0006] The control system includes a data acquisition module, a controlled module, a self-control module, and an execution module;

[0007] The acquisition module continuously sends a high-level signal y when urine is present. h The automatic control module collects analog sound signals x. raw The audio digital signal sequence y is generated through filtering and analog-to-digital conversion and transmitted to the controlled module, which continuously acquires the video digital signal sequence Y. vid And transmit it to the automatic control module and the execution module;

[0008] The controlled module senses user touch or receives remote control signals to generate control commands. It converts the sound digital signal sequence y into Mel-frequency cepstral coefficients MF. Based on a progressive binary classification model, it identifies whether the sound digital signal sequence y is an infant's cry or human language. If it is an infant's cry, it generates a prompt command and sends it to the execution module. If it is human language, it uses a deep hidden Markov model combined with a context-constrained word segmentation model to convert the Mel-frequency cepstral coefficients MF into speech content z and splits and recombines it into a speech phrase sequence Ph based on semantic dependencies. Using a skip-word model, it converts the speech phrases in the trigger phrase, functional phrase, and speech phrase sequence Ph into vector representations and performs matching verification. When the trigger phrase is successfully matched and at least one functional phrase is successfully matched, a control command is generated based on the matched functional phrase. In other cases, the sound digital signal sequence y is ignored. After the control command is generated, it is transmitted to the execution module.

[0009] The automatic control module periodically extracts video digital signal sequences Y. vid And convert it into a time segment video. The system identifies the infant's sleep state through facial recognition and optical flow detection, generates control or prompt instructions based on a rule engine, and transmits them to the execution module. A high-level signal (y) is then activated. h During the receiving period, prompt instructions are periodically sent to the execution module;

[0010] The execution module executes control commands and receives the video digital signal sequence Y. vid It also uploads data to the cloud platform, receives prompts and sends them to all paired devices via the cloud platform. Paired devices are mobile devices that are registered on the cloud platform and bound to the control system.

[0011] Furthermore, the data acquisition module includes a urine sensing unit, a sound sensing unit, and a video monitoring unit;

[0012] The urine sensing unit senses urine contact based on the principle of capacitive sensing and the conduction characteristics of an NMOS transistor, generating a high-level signal y when urine comes into contact with it. h The signal is then transmitted to the automatic control module, which stops generating a high-level signal y when urine is cleared. h ;

[0013] The sound sensing unit collects analog sound signals x raw The signal is converted into a digital audio signal sequence y by passing through a filter bank, an analog-to-digital converter, and an adaptive filtering algorithm, and then transmitted to the controlled module.

[0014] The video surveillance unit collects ambient light intensity in real time and adaptively determines the supplementary lighting brightness, continuously shooting to generate surveillance video X. vid It is then converted into a video digital signal sequence Y using the H.265 encoding algorithm. vid Synchronously send video digital signal sequence Y vid To the controlled module and the execution module.

[0015] Furthermore, the urine sensing unit, based on the principle of capacitive sensing and the conduction characteristics of NMOS transistors, senses urine contact through the following specific steps:

[0016] Connect the capacitor to the bias circuit, and connect the gate and source of the NMOS transistor to the bias circuit and ground, respectively. The bias circuit includes a resistor R and a DC power supply V. cc ;

[0017] When there is no urine contact, the normal current I0 of the bias circuit is equal to the DC power supply V. cc Divide by resistance R and normal capacitive reactance The sum of these, normal tolerance and resistance The voltage is inversely proportional to the normal capacitance C0, which is the capacitance when the capacitor dielectric is air. This is the normal gate-source voltage of the NMOS transistor. DC power supply V ccAdd normal resistance The product of the voltage and the normal current I0 is less than or equal to the breakdown voltage V. th The NMOS transistor is in the off state;

[0018] When urine comes into contact with the capacitor, the capacitor dielectric becomes urine, and the capacitor's non-state capacitive reactance... Less than normal capacitance The abnormal current I1 of the bias circuit is greater than the normal current I0, and the abnormal gate-source voltage of the NMOS transistor is... Equal to DC power supply V cc Adding non-normal capacitive The product of the abnormal current I1 and the breakdown voltage V is greater than the breakdown voltage V. th The NMOS transistor is in the on state, generating a drain-source current I. ds ;

[0019] The drain-source current I is obtained by sampling the resistor. ds Converted into output signal voltage V out It is connected to a comparator circuit, and the comparator outputs a signal voltage V. out Greater than the reference voltage V ref A high-level signal y is generated at this time. h The signal is transmitted to the automatic control module; otherwise, the output signal voltage V is ignored. out ;

[0020] When the urine is cleared, the capacitor's dielectric returns to air, the NMOS transistor turns off, and the high-level signal y stops being output. h .

[0021] Furthermore, the process of generating a digital audio signal sequence y through filter banks, analog-to-digital converters, and adaptive filtering algorithms includes the following specific steps:

[0022] The analog sound signal x raw Input filter bank, for frequencies less than the minimum human voice frequency f min and greater than the maximum human voice frequency f max The frequency range is filtered multiple times to generate a filtered analog signal x. hp ;

[0023] The analog signal x is filtered using an analog-to-digital converter. hp Discrete sampling is performed to generate an initial digital signal sequence y of sound. ori ;

[0024] Initialize the initial digital signal y of the sound at the first sampling time. ori (1) The weight coefficient vector w(1) and the positive definite symmetric matrix P(1);

[0025] A reference noise signal sequence d is generated by discretization based on a priori white noise model. ref , where d ref (n) is the reference noise signal at the nth sampling time, where n = 1, ..., N, and N is the total number of sampling times;

[0026] Based on the initial digital signal y of the sound at the nth sampling time ori The initial digital signal of sound at the first M-1 sampling times (n) is used to construct the initial input vector y of sound at the nth sampling time. ori (n), if there are less than M-1 sampling times before the nth sampling time, the insufficient part is padded with 0;

[0027] The weight coefficient vector w(n) at the nth sampling time is compared with the initial sound input vector y. ori The inner product of (n) is taken as the audio digital signal y(n) at the nth sampling time;

[0028] The error signal e(n) at the nth sampling time is used to reflect the effect of adaptive filtering, and is equal to the audio digital signal y(n) plus the reference noise signal d. ref (n) Subtract the initial digital signal y from the sound ori (n);

[0029] The gain vector g(n) at the nth sampling time is calculated based on the forgetting factor λ and the positive definite symmetric matrix P(n-1) at the (n-1)th sampling time. The weight coefficient vector w(n) at the nth sampling time is adjusted in combination with the error signal e(n) to generate the weight coefficient vector w(n+1) at the (n+1)th sampling time.

[0030] Based on the forgetting factor λ and the gain vector g(n) at the nth sampling time, update the positive definite symmetric matrix P(n-1) at the (n-1)th sampling time to obtain the positive definite symmetric matrix P(n) at the nth sampling time.

[0031] The algorithm minimizes the weighted sum of squares J of the error signal through multiple iterations. When the algorithm converges, it generates a sequence of digital audio signals y. In each iteration, the weight coefficient vector and positive definite symmetric matrix at N sampling times are updated.

[0032] Furthermore, the controlled modules include a touch control unit, a remote control unit, and a voice control unit;

[0033] The touch control unit generates a touch distribution network based on parasitic capacitance sensing, verifies the overlap rate with the button layout network to make decisions on generating control commands, providing vibration feedback and sending them to the execution module;

[0034] The remote control unit receives remote control signals, generates corresponding control commands, and forwards them to the execution module.

[0035] The voice control unit filters and logarithmically transforms the audio digital signal sequence y in the frequency domain with a Mel filter bank to generate Mel frequency cepstral coefficients MF in the cepstral domain. A progressive binary classification model is used to sequentially determine whether the audio digital signal sequence y is an infant cry or human speech based on the Mel frequency cepstral coefficients MF. If it is an infant cry, a prompt command is generated and sent to the execution module. If it is human speech, a deep hidden Markov model is used to transform the Mel frequency cepstral coefficients MF into speech content z. A context-constrained word segmentation model is used to split the speech content z into a word segmentation sequence Wo based on semantic dependencies. A speech phrase sequence Ph is constructed by combining these segments using a sliding window. A skip-word model is used to match and verify the trigger phrase, functional phrase, and speech phrases in the speech phrase sequence Ph. If the trigger phrase and at least one functional phrase match the speech phrase sequence Ph, the corresponding control command is generated based on the matched functional phrase; otherwise, the audio digital signal sequence y is ignored.

[0036] Furthermore, the process of generating a touch distribution network based on parasitic capacitance sensing and verifying its overlap rate with the button layout network includes the following specific steps:

[0037] The peak value of local charge change caused by parasitic capacitance formed by coupling when the user makes contact is obtained and it is determined whether it is greater than the effective peak value of the touch.

[0038] If the local charge peak value is less than the effective touch peak value, ignore the local charge change. If the local charge peak value is greater than or equal to the effective touch peak value, the grid in the touch distribution network that overlaps with the local charge change area is recorded as the active grid.

[0039] Call the button layout net and verify its overlap with the touch distribution net. Confirm the overlap rate between the grid area of ​​each button number in the button layout net and the active grid in the touch distribution net, and determine whether the maximum overlap rate is greater than or equal to the overlap rate threshold.

[0040] If the maximum overlap rate is less than the overlap rate threshold, ignore the touch. If the maximum overlap rate is greater than or equal to the overlap rate threshold, generate the corresponding control command based on the button number with the maximum overlap rate.

[0041] Furthermore, the process of filtering and performing logarithmic operations on the audio digital signal sequence y in the frequency domain with a Mel filter bank to convert it to the cepstral domain and generate Mel frequency cepstral coefficients MF includes the following specific steps:

[0042] The audio digital signal sequence y is processed by Discrete Fourier Transform to generate a spectrum set F. The spectrum set F includes the spectrum sequences at N sampling times, and the spectrum sequence f at the nth sampling time. n Includes frequency domain signals at K frequencies, where N and K are the total number of sampling times and the total number of frequencies, respectively.

[0043] Based on the Mel transform equation, the frequency range of the audio digital signal sequence y is transformed into the Mel frequency range, and the Mel frequency range is further divided into Q Mel frequency segments. The center Mel frequency of each Mel frequency segment is determined, where Q is the total number of Mel frequency segments.

[0044] Based on the center Mel frequency of the (q-1)th Mel frequency band The center Mel frequency of the qth Mel frequency band and the center Mel frequency of the (q+1)th Mel frequency band Set up a triangular filter for the q-th Mel frequency band, where q = 1, ..., Q;

[0045] The spectral sequences at N sampling times in the spectral set F are multiplied by Q triangular filters and the logarithm is taken to generate a logarithmic domain signal set LF. The logarithmic domain signal set LF includes the logarithmic domain signal sequences at N sampling times, and the logarithmic domain signal sequence at the nth sampling time is lf. n Includes Q logarithmic domain signals, the logarithmic domain signal sequence lf n The q-th logarithmic domain signal lf n (q) is the spectral sequence f n The sum of the products of the frequency domain signal at each frequency and the frequency response of the triangular filter in the q-th Mel frequency band at that frequency;

[0046] The logarithmic domain signal set LF is transformed to the cepstral domain by discrete cosine transform, generating Mel frequency cepstral coefficients MF, which include Mel frequency cepstral eigenvectors at N sampling times.

[0047] Specifically, the progressive binary classification model includes a first support vector machine and a second support vector machine. The first support vector machine classifies the Mel frequency cepstral coefficients (MF) into either infant crying or the first residue class in the high-dimensional space. When the classification result of the first support vector machine is the first residue class, the second support vector machine further classifies the Mel frequency cepstral coefficients (MF) into either human language or the second residue class in the high-dimensional space. The first residue class is equal to the union of human language and the second residue class, and infant crying is mutually exclusive with the first residue class, while human language is mutually exclusive with the second residue class.

[0048] Specifically, deep hidden Markov models include deep networks and hidden Markov models;

[0049] The input to the deep network is the Mel frequency cepstral coefficients (MF). The Mel frequency cepstral coefficients (MF) are linearly modulated multiple times through a linear layer to adjust to a certain dimension. The posterior probability distribution of the state vector is generated through the Softmax function. The posterior probability distribution records the posterior probability of each state in the state vector.

[0050] Hidden Markov Models (HMMs) define state vectors based on phonemes. These state vectors include silence, phonation initiation, stable phonation, and phonation termination states. A state transition probability matrix is ​​defined to describe the probability of transitioning from one state to another. The observation probability distribution is defined as the posterior probability distribution of the state vector, reflecting the probability of observing the Mel-frequency cepstral coefficient (MF) in a single state. An initial state probability distribution is defined for the initial time step. At each time step, the probability of transitioning from the current state to the next state and the probability of observing the MF are determined based on the state transition probability matrix and the observation probability distribution. The Viterbi algorithm is used to calculate the maximum probability path from the current state to the remaining states and the path information is recorded. The optimal state sequence is obtained by backtracking. Based on the pronunciation dictionary, the phonemes in the optimal state sequence are combined to generate the speech content z.

[0051] Specifically, the context-constrained word segmentation model is used to split speech content z into word segmentation sequence Wo. The context-constrained word segmentation model includes a bidirectional long short-term memory network layer and a conditional random field.

[0052] The bidirectional long short-term memory network layer includes forward LSTM and backward LSTM. Forward LSTM and backward LSTM process speech content z in forward and backward order, respectively. Through input gate and forget gate, the influence between adjacent Chinese characters is learned, capturing the forward and backward semantic associations of speech content z, and generating semantic features.

[0053] Conditional random fields (CRFs) calculate the probability that the feature value of each dimension of the semantic features belongs to a single part-of-speech tag based on the state feature function to obtain all candidate word segmentation sequences. By focusing on the transition relationship between part-of-speech tags through the transition feature function, the candidate word segmentation sequences are constrained. The conditional probability of each candidate word segmentation sequence is calculated. The optimal part-of-speech tag output order is found through the Viterbi algorithm, and the speech content z is labeled for word segmentation to generate the word segmentation sequence Wo.

[0054] Specifically, the construction of the speech phrase sequence Ph relies on a sliding window. The width of the sliding window is the maximum phrase length among the trigger phrase and all functional phrases. The starting edge of the sliding window is placed at the first segment in the segmentation sequence Wo. The segments within the width of the sliding window are combined into speech phrases. Then, the sliding window is shifted backward by one segment, and new speech phrases are generated again. This process continues until the ending edge of the sliding window reaches the last segment, at which point the last speech phrase is generated. All speech phrases are then used to construct the speech phrase sequence Ph according to their generation order.

[0055] Furthermore, the skip word model is used to convert the speech phrases in the trigger phrase, functional phrase, and speech phrase sequence Ph into vector representations and perform matching verification, including the following specific steps:

[0056] The skip-word model extracts the central word of the input phrase and determines the co-occurrence probability of the central word and the context words. Based on the corpus, the central word is converted into one-hot encoding. The probability distribution of the input phrase is generated through two linear layers and a Softmax function. Backpropagation is used to adjust the weights of the linear layers to minimize the error between the probability distribution and the co-occurrence probability of the input phrase. The weights of the first linear layer are used as the vector representation of the input phrase. The input phrase includes trigger phrases, function phrases, and speech phrases in the speech phrase sequence Ph. The vector representation includes trigger vector representation, function vector representation, and speech vector representation.

[0057] Calculate the trigger cosine similarity between the trigger vector representation and each speech vector representation, and simultaneously calculate the functional cosine similarity between each function vector representation and each speech vector representation, and compare them sequentially with the similarity threshold.

[0058] If both trigger cosine similarity and function cosine similarity are greater than or equal to the similarity threshold, corresponding control instructions are generated based on all function phrases whose function cosine similarity is greater than or equal to the similarity threshold.

[0059] If there is no trigger cosine similarity or functional cosine similarity greater than or equal to the similarity threshold, ignore the audio digital signal sequence y.

[0060] Furthermore, the self-control module includes a bedwetting warning unit and a video monitoring unit;

[0061] The bedwetting indicator unit receives a high-level signal y h It records the reception time, periodically generates and sends prompt commands to the execution module, until no high-level signal y is received. h Clear the receiving time;

[0062] The video monitoring unit periodically captures the video digital signal sequence Y after the control system is activated. vid The video segment is then converted and generated based on the H.265 encoding algorithm. Video analysis of time periods using facial recognition and optical flow detection The system identifies the infant's sleep state, generates control or prompt instructions based on the rule engine, and transmits them to the execution module. The infant's sleep state includes deep sleep, light sleep, semi-awake, awake, and abnormal sleep.

[0063] Furthermore, video over time can be analyzed using facial recognition and optical flow detection. Identifying an infant's sleep state involves the following specific steps:

[0064] Sampling period video Construct a sequence of sampled frame images from T sampled frames.

[0065] The sampled frame image sequence was located using Dlib's face detector and a 68-facial landmark model. The coordinates of the baby's face region and key points around the eyes in each sampled frame image are calculated. The vertical distance between the midpoint of the upper eyelid and the midpoint of the lower eyelid is calculated based on the coordinates of the key points around the eyes in each sampled frame image and the average value is taken. The average vertical distance is obtained and divided by the maximum vertical distance to generate the eye-opening ratio.

[0066] Extracting sampled frame image sequences from the initial convolutional layer of the MobileNetV3 model. For each sampled frame image, the feature map is concatenated using a spatial attention mechanism, combining the average pooling result and the global pooling result of the channel sub-maps of the feature map in multiple channels. The spatial attention map of each sampled frame image is then generated through convolution and the Sigmoid function.

[0067] The spatial attention map is connected to the depthwise separable convolutional layer of the MobileNetV3 model. Dimensionality upscaling, feature extraction, and dimensionality reduction of the spatial attention map are achieved through pointwise convolution, depthwise convolution, and pointwise convolution. After global average pooling, fully connected layers, and the Softmax function, the probability distribution of the expression level of each sampled frame image is generated, and the average probability distribution of the expression level is calculated. The expression level includes calm, moderate, and rich.

[0068] Detection of sampled frame image sequences using optical flow method The temporal changes of pixels are analyzed, and a luminance identity is defined to describe the constant brightness of corresponding pixels between adjacent frames. Based on the assumption of spatiotemporal continuity, the offset of corresponding pixels between adjacent frames is small. The image sequence describing the sampled frames is found through TV-L1 iterative optimization. The optical flow feature sequence, in which the optical flow feature map describes the movement trajectory and velocity of pixels between adjacent frames over time;

[0069] The optical flow feature sequence is input into the temporal convolutional layer. The temporal convolutional layer captures the temporal changes of the optical flow features through two identical residual blocks. Then, the fully connected layer reduces the dimension and outputs an aggregated optical flow feature map with a dimension equal to that of the optical flow feature map. The residual block includes dilated convolution and residual structure. The dilated convolution inserts zero elements between the convolution kernels to expand the receptive field. The residual structure superimposes the output of the dilated convolution with the input.

[0070] The aggregated optical flow feature map is reduced in dimension and mapped by a fully connected layer and a softmax function to generate a probability distribution of action amplitude, which includes small, medium and large action amplitudes.

[0071] Based on the scores for each level of facial expression intensity and movement amplitude, the facial expression intensity probability distribution and movement amplitude probability distribution are weighted and summed to generate facial expression score and movement score. Based on the reference weights, the eye-opening ratio, facial expression score and movement score are weighted and summed again to generate a comprehensive sleep score. The infant's sleep state is determined based on the comprehensive sleep score range. Infant sleep state includes deep sleep, light sleep, semi-awake, awake and abnormal.

[0072] Specifically, the rule engine establishes a correspondence between the baby's sleep state and the target operating state of the variable speed synchronous motor and the speaker. After determining the baby's sleep state, it generates corresponding control commands or prompts based on the current operating state of the variable speed synchronous motor and the speaker, so that the variable speed synchronous motor and the speaker reach the target operating state. The control commands include power-on command, power-off command, swing start command, swing stop command, swing acceleration command, swing deceleration command, music playback command, music stop command, story playback command, and story stop command.

[0073] Furthermore, the execution module includes a power switching unit, a variable speed synchronous motor execution unit, a sound playback unit, and a cloud communication unit;

[0074] The power switching unit receives power-on or power-off commands and turns the circuit switch on or off.

[0075] The variable speed synchronous motor actuator receives the swing start command or swing stop command, starts the variable speed synchronous motor according to the previous motor gear or records the current motor gear and stops the variable speed synchronous motor, receives the swing acceleration command or swing deceleration command, and adjusts up or down one gear within the allowed motor gear range. The motor gears include low gear, medium gear and high gear.

[0076] The sound playback unit receives music playback instructions or story playback instructions, retrieves the previously played music or story, plays the next piece of music or the next story through the speaker, receives music stop instructions or story stop instructions, records the currently playing music or story, and turns off the speaker;

[0077] The cloud communication unit receives the video digital signal sequence Y. vid The system sends prompts and instructions to the cloud platform, which is used to establish remote communication between the control system and the paired device.

[0078] The controlled device, controlled by the control system, includes a cradle body, a base, a drive unit, a universal ball bearing, a universal rocking plate, a cradle body connecting flange, a limit rubber belt, a urine sensor, and an integrated processing device.

[0079] The cradle is used to accommodate a baby while sleeping;

[0080] The base is used to increase stability;

[0081] The drive unit includes a variable speed synchronous motor and a variable speed synchronous motor rocker arm. When the variable speed synchronous motor starts, it drives the variable speed synchronous motor rocker arm to swing. The variable speed synchronous motor rocker arm is fixedly connected to the universal swing disk.

[0082] Universal ball bearings are used to support universal sway discs;

[0083] The universal swing plate and the rocker arm of the variable speed synchronous motor maintain the same swing direction.

[0084] The cradle body connecting flange ensures that the cradle body and the universal rocking plate swing in the same direction.

[0085] The limiting rubber belt prevents the universal swivel plate from derailing;

[0086] The urine sensor is used to detect urine contact and generate a high-level signal;

[0087] The integrated processing unit is used to acquire audio digital signal sequences and monitor video digital signal sequences. It provides four ways for users to control the opening and closing of the drive device and the adjustment of gears, as well as the playback and stopping of music and stories, through touch control, remote control, voice control, and automatic control. It also sends video digital signal sequences and prompts to the cloud platform.

[0088] Furthermore, the integrated processing device includes a keypad screen, a rear camera, a microphone array, a speaker, a wireless communication module, and a processor;

[0089] The keypad displays buttons for users to touch.

[0090] The rear camera monitors and records the cradle body;

[0091] Microphone arrays are used to collect analog sound signals and convert them into digital sound signal sequences;

[0092] The speaker is controlled by the processor to play or stop playing music or a story.

[0093] The wireless communication module receives remote control signals and forwards them to the processor; it also receives video digital signal sequences and prompts and uploads them to the cloud platform.

[0094] During high-level signal reception, the processor generates prompt instructions, recognizes user touch buttons to generate control instructions, receives remote control signals to generate control instructions, receives and determines whether the audio digital signal sequence is an infant's cry or human speech. If it is an infant's cry, a prompt instruction is generated; if it is human speech, it is converted into speech content and matched and verified to decide whether to generate a control instruction. The processor periodically captures video digital signal sequences and transcodes them into time-segment video to identify the infant's sleep state. Based on the rule engine, control instructions or prompt instructions are generated. Control instructions are used to control the drive device and speaker, while prompt instructions are handed over to the wireless communication module.

[0095] Compared with existing technologies, this invention uses the principle of capacitive sensing and the conduction characteristics of NMOS transistors to sense urine contact and generate a high-level signal. During the high-level signal generation period, it continuously generates prompts to automatically detect bedwetting in infants. Building upon traditional button touch, it establishes a cloud platform to bind the controlled device and paired equipment, enabling remote control based on remote control signals. The invention converts the acquired digital sound signal sequence into Mel-frequency cepstral coefficients. Based on the recognition results of a progressive binary classification model, it generates prompts for infant crying or human language selection, or uses speech processing technologies such as deep hidden Markov models, context-constrained word segmentation models, and skip-word models to analyze speech content and generate corresponding control commands, achieving automatic detection of infant crying and effective voice control. It periodically captures video digital signals and identifies the infant's sleep state through transcoding, facial recognition, and optical flow detection. Based on a rule engine, it automatically generates control commands, achieving adaptive intelligent control based on the infant's sleep state. This integrated control method combining touch, remote, voice, and automation improves the intelligence level of the infant cradle. Attached Figure Description

[0096] Figure 1 This is a schematic diagram of the control system in this invention;

[0097] Figure 2 This is a schematic diagram of overlap rate verification in this invention;

[0098] Figure 3 This is a schematic diagram illustrating how facial recognition and optical flow detection are used to identify an infant's sleep state in this invention.

[0099] Figure 4 This is a structural diagram of the controlled device in this invention.

[0100] Reference numerals: 1. Cradle body; 2. Base; 3. Drive device; 31. Variable speed synchronous motor; 32. Motor rocker arm; 4. Universal ball bearing; 5. Universal rocker plate; 6. Cradle body connecting flange; 7. Limiting rubber belt; 8. Urine sensor; 9. Integrated processing device. Detailed Implementation

[0101] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0102] Example 1

[0103] like Figure 1 As shown, the present invention discloses a control system, including a data acquisition module, a controlled module, a self-control module, and an execution module;

[0104] The acquisition module continuously sends a high-level signal y while urine is present. hThe automatic control module collects analog sound signals x. raw After filtering and analog-to-digital conversion, an audio digital signal sequence y is generated and transmitted to the controlled module, which continuously monitors and acquires the video digital signal sequence Y. vid And transmit it to the automatic control module and the execution module;

[0105] The controlled module senses user touch, identifies touch buttons to generate control commands, receives remote control signals from the cloud platform and converts them into control commands, receives the audio digital signal sequence y and converts it into Mel-frequency cepstral coefficients MF. Using a progressive binary classification model, it analyzes the Mel-frequency cepstral coefficients MF twice via support vector machines to sequentially identify whether the audio digital signal sequence y is an infant cry or human speech. If it is an infant cry, a prompt command is generated and sent to the execution module. If it is human speech, a deep hidden Markov model is used to convert the Mel-frequency cepstral coefficients MF into speech content z. A context-constrained word segmentation model splits and reorganizes the speech content z into a speech phrase sequence Ph based on the semantic dependencies of the speech content z. A skip-word model is used to separate each speech phrase in the trigger phrase, each functional phrase, and each speech phrase in the speech phrase sequence Ph. The speech is converted into a vector representation and matched based on similarity. If the trigger phrase matches the speech phrase sequence Ph and at least one functional phrase matches the speech phrase sequence Ph, the corresponding control command is generated by searching the instruction encoding table based on the matched functional phrase. If the trigger phrase fails to match or all functional phrases fail to match, the sound digital signal sequence y is ignored. After the control command is generated, it is transmitted to the execution module. The control command is a binary code that uniquely identifies the corresponding function, including power on command, power off command, swing start command, swing stop command, swing acceleration command, swing deceleration command, music play command, music stop command, story play command, and story stop command. The instruction encoding table stores the functional phrase, button number, remote control signal, and control command corresponding to each function.

[0106] The automatic control module responds to the high-level signal y. h During reception, prompts are periodically sent to the execution module, nodes are periodically generated, and at each node, a video digital signal sequence Y is captured. vid The video digital signal from the previous node to the current node is converted into a time segment video. The system identifies the infant's sleep state through facial recognition and optical flow detection, generates control or prompt instructions based on the rule engine, and transmits them to the execution module.

[0107] The execution module executes control commands and receives the video digital signal sequence Y. vidThe system simultaneously uploads the video data to the cloud platform's database, receives prompts and sends them to all paired devices via the cloud platform. The paired devices are mobile devices that the user has registered on the cloud platform and bound to the control system. The control system can pair multiple successfully registered mobile devices simultaneously.

[0108] Furthermore, the data acquisition module includes a urine sensing unit, a sound sensing unit, and a video monitoring unit;

[0109] The urine sensing unit detects urine contact based on the principle of capacitive sensing and the conduction characteristics of an NMOS transistor. When urine comes into contact with the urine sensor, the medium between the capacitor electrodes changes from air to urine. This increases the current in the bias circuit, causing the NMOS transistor to conduct and generating a high-level signal y. h The signal is transmitted to the automatic control module. When the urine is cleared, the NMOS transistor turns to the off state, and a high-level signal y is generated. h Stop generating;

[0110] The sound sensing unit collects analog sound signals x through a microphone array. raw The analog sound signal is initially filtered using a filter bank, and after passing through an analog-to-digital converter, the residual noise is further suppressed by an adaptive filtering algorithm to generate a digital sound signal sequence y and transmit it to the controlled module.

[0111] The video monitoring unit uses a digital optical sensor to collect ambient light intensity in real time, determine the light intensity range of the ambient light, and determine the corresponding brightness of the infrared supplementary light. A high-definition camera continuously captures the baby's state in the cradle, generating monitoring video X. vid The H.265 encoding algorithm is used to encode the surveillance video X. vid Converted to video digital signal sequence Y vid The light intensity range is pre-divided, and each light intensity range is calibrated according to the corresponding supplementary light brightness. The wavelength of the infrared supplementary light is 850nm, which is invisible to the human eye and will not affect the baby's sleep. The H.265 encoding algorithm is an existing video encoding and transmission algorithm.

[0112] Furthermore, the urine sensing unit, based on the principle of capacitive sensing and the conduction characteristics of NMOS transistors, senses urine contact through the following specific steps:

[0113] Connect the capacitor to the bias circuit to form a complete circuit loop. Connect the gate of the NMOS transistor to the bias circuit, and ground the source. The bias circuit includes a resistor R and a DC power supply V. cc ;

[0114] When there is no urine contact, the dielectric between the capacitor electrodes is air, the capacitance is the normal capacitance C0, and the normal current I0 of the bias circuit should be equal to the DC power supply V.cc Divide by resistance R and normal capacitive reactance The sum is as follows:

[0115]

[0116] Among them, normal resistance It is inversely proportional to the normal capacitance C0, as follows:

[0117]

[0118] Among them, f cc The operating frequency of the power supply;

[0119] Normal gate-source voltage of NMOS transistor DC power supply V cc Add capacitor normal voltage Specifically: Normal gate-source voltage Less than or equal to the breakdown voltage V th The source and drain of the NMOS transistor are in the off state;

[0120] When urine comes into contact with the capacitor, the dielectric between the capacitor electrodes becomes urine, and the capacitance becomes an abnormal capacitance C1. Since the capacitance of a capacitor is equal to the product of the electrode plate area and the dielectric constant divided by the distance between the electrode plates, and the dielectric constant of urine is greater than that of air, the abnormal capacitance C1 is greater than the normal capacitance C0. This results in an abnormal capacitive reactance. Less than normal capacitance The abnormal current I1 is greater than the normal current I0, and the abnormal voltage of the capacitor is... Greater than the normal voltage of the capacitor The abnormal gate-source voltage of an NMOS transistor Greater than the breakdown voltage V th When the source and drain of an NMOS transistor are in the ON state, a drain-source current I is generated. ds ;

[0121] The drain-source current I is obtained by sampling the resistor. ds Converted into output signal voltage V out It is then connected to a comparator circuit, which outputs a signal voltage V. out With the preset reference voltage V ref Compare, if the output signal voltage V out Greater than the reference voltage V ref , generates a high-level signal y h And transmit it to the automatic control module, if the output signal voltage V out Less than or equal to the reference voltage V ref Then ignore the output signal voltage V outCompared to directly judging based on the level signal generated after the NMOS transistor is turned on, the comparator circuit effectively avoids misjudgments caused by noise interference or environmental electromagnetic interference. Furthermore, the level signal generated after the NMOS transistor is turned on does not conform to the logic level standard and still requires additional circuitry for conversion before transmission. The high-level signal y output by the comparator... h It conforms to logic level standards and is highly compatible with any chip or circuit architecture.

[0122] When the urine is completely removed, the capacitor's dielectric gradually returns to air, and the capacitance gradually returns from the abnormal state capacitance C1 to the normal state capacitance C0. The drain-source current I of the NMOS transistor... ds Gradually decrease to 0, when the output signal voltage V out Less than or equal to the reference voltage V ref When the comparator stops outputting the high-level signal y, h .

[0123] Furthermore, the process of using filter banks, analog-to-digital converters, and adaptive filtering algorithms to filter, convert, and perform secondary filtering on the analog sound signal to generate a digital sound signal sequence y includes the following specific steps:

[0124] The analog sound signal x raw The input filter bank includes a bandpass filter, a low-pass filter, and a high-pass filter. The bandpass filter allows [f]... min ,f max Signals within the frequency range pass through, where f min and f max These are the minimum and maximum human voice frequencies, respectively. Low-pass and high-pass filters are used to filter frequencies less than the minimum human voice frequency f. min and greater than the maximum human voice frequency f max The frequency range is further filtered to generate a filtered analog signal x. hp In this embodiment, considering the subsequent need to recognize infant cries or human speech, the minimum human voice frequency f is... min and the maximum human voice frequency f max They are 80Hz and 8kHz respectively;

[0125] The analog signal x is filtered using an analog-to-digital converter. hp Discrete sampling is performed in time to generate an initial digital signal sequence y of the sound. ori ={y ori (1),…,y ori (n),…,y ori (N)},y ori (n) represents the initial digital signal of the sound at the nth sampling time, where n = 1, ..., N, N is the total number of sampling times, and the sampling frequency is f. sSatisfying the Nyquist sampling theorem, i.e., the sampling frequency f s Greater than the maximum human voice frequency f max 2 times;

[0126] Initialize and generate the initial digital signal y of the sound at the first sampling time. ori (1) The weight coefficient vector w(1) = [w1(1),...,w m (1),…,w M (1)] and positive definite symmetric matrix P(1)=δ -1 ·I, where w m (1) is the initial digital signal y of the sound at the first sampling time. ori (1) The m-th weight coefficient, m=1,…,M, δ is a minimal positive number, I is the identity matrix with dimension M×M, and M is the reference number;

[0127] The reference noise signal sequence d is generated by discretizing a prior white noise model that follows a Gaussian distribution. ref ={d ref (1),…,d ref (n),…,d ref (N)}, where d ref (n) is the reference noise signal at the nth sampling time, where n = 1, ..., N, and N is the total number of sampling times;

[0128] The initial digital signal y of the sound at the nth sampling time is obtained based on the reference number M. ori (n) The initial digital signal y of the sound from the previous M-1 sampling times. ori (n-1),…,y ori (n-M+1), construct the initial sound input vector y at the nth sampling time. ori (n)=[y ori (n),…,y ori [(n-M+1)], where if n < M-1, meaning the initial digital signal of the sound at M-1 sampling times cannot be obtained before the nth sampling time, then the insufficient part is padded with 0. Taking n=2 and M=3 as an example, the initial sound input vector y ori (2)=[y ori (2),y ori [(1),0];

[0129] The weight coefficient vector w(n) = [w1(n),...,w] at the nth sampling time is... m (n),…,w M [n] and the initial sound input vector y ori The inner product of (n) is taken as the audio digital signal y(n) at the nth sampling time after adaptive filtering, and the specific formula is:

[0130] Calculate the error signal e(n) = y(n) + d at the nth sampling time. ref (n)-y ori The error signal e(n) reflects the effect of adaptive filtering. The smaller the error signal e(n), the better the adaptive filtering will produce the initial digital signal y at the nth sampling time. ori The more complete the noise filtering in (n);

[0131] The gain vector g(n) at the nth sampling time is calculated based on the forgetting factor λ and the positive definite symmetric matrix P(n-1) at the (n-1)th sampling time. The specific formula is as follows:

[0132]

[0133] The weight coefficient vector w(n) at the nth sampling time is adjusted based on the gain vector g(n) and the error signal e(n) at the nth sampling time to generate the weight coefficient vector w(n+1) at the (n+1)th sampling time. Specifically, w(n+1) = w(n) + g(n)e(n), that is, the larger the error signal e(n), the larger the adjustment amplitude.

[0134] To ensure convergence, the positive definite symmetric matrix P(n) at the nth sampling time is updated, and the specific formula is as follows:

[0135]

[0136] The ultimate goal of the adaptive filtering algorithm is to minimize the weighted sum of squares of the error signal through multiple iterations. When the algorithm converges, it generates a sequence of digital audio signals y, in which the weight coefficient vector and positive definite symmetric matrix at N sampling times are updated sequentially in each iteration.

[0137] Furthermore, the controlled modules include a touch control unit, a remote control unit, and a voice control unit;

[0138] The touch control unit is based on a capacitive touch sensor. It senses the user's touch position based on the principle of capacitive parasitics, generates a touch distribution network, and verifies the overlap rate with the button layout network. Only when the overlap rate between the active grid in the touch distribution network and the grid area of ​​a single button is greater than or equal to the overlap rate threshold, the instruction encoding table is retrieved based on the button number, the corresponding control instruction is generated and sent to the execution module, and vibration feedback is provided to indicate that the touch is valid. In all other cases, the touch is ignored.

[0139] The remote control unit wirelessly receives remote control signals from the cloud platform, retrieves the instruction encoding table, generates control instructions that match the remote control signals, and forwards them to the execution module. The remote control signals are sent by the user through the pairing device. The cloud platform receives the remote control signals sent by the pairing device, determines the paired control system based on the registration information, and issues the commands.

[0140] The voice control unit receives the audio digital signal sequence y, converts it to the frequency domain, constructs a Mel filter bank for filtering, performs logarithmic operations, and converts it to the cepstral domain to generate Mel frequency cepstral coefficients MF. A progressive binary classification model is used to analyze the Mel frequency cepstral coefficients MF. A first support vector machine determines whether the audio digital signal sequence y is an infant cry. If it is an infant cry, a prompt command is generated and sent to the execution module. If it is not an infant cry, a second support vector machine determines whether the audio digital signal sequence y is human speech. If it is not human speech, the audio digital signal sequence y is ignored. If it is human speech, a deep hidden Markov model is used to convert the Mel frequency cepstral coefficients MF into speech. For content z, semantic features of speech content z are extracted and semantic dependencies are learned through a context-constrained word segmentation model. Speech content z is split into word segmentation sequence Wo. A sliding window is moved in the word segmentation sequence Wo to combine consecutive word segments within the sliding window into speech phrases, constructing a speech phrase sequence Ph. A skip-word model is used to convert the trigger phrase, each functional phrase, and each speech phrase in the speech phrase sequence Ph into corresponding vector representations and verify them by similarity matching. If the trigger phrase matches the speech phrase sequence Ph and at least one functional phrase matches the speech phrase sequence Ph, the corresponding control command is generated based on the command encoding table retrieved according to the matched functional phrase.

[0141] by Figure 2 For example, further, the process of generating a touch distribution network based on the principle of capacitive parasitics and verifying its overlap with the button layout network includes the following specific steps:

[0142] The system detects changes in local charge, acquires the peak value of local charge, and determines whether the peak value of local charge is greater than the effective peak value of touch. The change in local charge is caused by the user's touch. When the user touches the button area, it is coupled with the capacitive touch sensor below the button area. A parasitic capacitance is generated near the contact position in parallel with the parallel capacitor, which causes the amount of local charge to increase. The parallel capacitor is a stable capacitor that is uniformly distributed in the capacitive touch sensor when there is no touch.

[0143] If the local charge peak value is less than the effective touch peak value, it is considered an invalid touch, and the local charge change is ignored. If the local charge peak value is greater than or equal to the effective touch peak value, it is considered a valid touch, and the grids in the touch distribution network that overlap with the local charge change area are recorded as active grids. Figure 2The touch distribution grid is generated by dividing the button area into grids. When there is no user touch, all grids in the touch distribution grid are silent grids by default. When the user touches, the distribution of active grids in the touch distribution grid can effectively reflect the user's touch position in the button area.

[0144] Call the button layout net and verify its overlap with the touch distribution net. The button layout net uses a grid to divide the button area in the same way as the touch distribution net. For a single grid in the button layout net, if it is covered by a button, the grid value is the corresponding button number; if it is not covered by a button, the grid value is empty.

[0145] Confirm the overlap rate between the grid area of ​​each button number in the button layout network and the active grid in the touch distribution network, and determine whether the maximum overlap rate is greater than or equal to the overlap rate threshold. The overlap rate is equal to the number of grids in the grid area of ​​a single button number that overlap with the active grid divided by the total number of grids of a single button number.

[0146] If the maximum overlap rate is less than the overlap rate threshold, it means that the user's touch position does not overlap with any button, and is considered an invalid touch, which is ignored.

[0147] If the maximum overlap rate is greater than or equal to the overlap rate threshold, the instruction encoding table is retrieved based on the key number of the maximum overlap rate (in...). Figure 2 If the button number is 2, generate the corresponding control command.

[0148] Furthermore, the process of converting the audio digital signal sequence y to the frequency domain and constructing a Mel filter bank for filtering, followed by logarithmic operations and conversion to the cepstral domain to generate Mel frequency cepstral coefficients MF, includes the following specific steps:

[0149] The audio digital signal sequence y = {y(1), ..., y(n), ..., y(N)} is transformed to the frequency domain using Discrete Fourier Transform to generate a spectrum set F = {f1, ..., f2}. n ,…,f N}, where f n ={f n (0),…,f n (k),…,f n (K-1)} represents the audio digital signal y(n) at the nth sampling time, which is converted into the corresponding spectrum sequence by the Discrete Fourier Transform, f n (k) represents the k-th frequency f of the digital audio signal y(n). k The specific calculation formula for the frequency domain signal is as follows:

[0150]

[0151] j is an imaginary number, j 2=-1, y(n) l Let y(n) be the frame signal of the audio digital signal y(n) in the l-th frame, l = 0, ..., L-1, where L is the frame length of each audio digital signal, k is the frequency index, k = 0, ..., K-1, where K is the total number of frequencies, and the discrete Fourier transform has symmetry, that is, the total number of frequencies K in the frequency domain is equal to the frame length L in the time domain.

[0152] Since the human auditory system perceives sound frequencies approximately on a Mel scale, given a digital sound signal sequence y filtered by a filter bank, its frequency range is [f]. min ,f max ], f min and f max These are the minimum and maximum human voice frequencies, based on the Mel frequency f. m The Mel frequency range is determined by the Mel transform equation of frequency f. and The minimum Mel frequency and the maximum Mel frequency are respectively, and the Mel transform equation is as follows:

[0153]

[0154] Mel frequency range Divide into Q Mel frequency bands and determine the center Mel frequency of each Mel frequency band. Let q be the center Mel frequency of the qth Mel frequency band, where q = 1, ..., Q, and Q is the total number of Mel frequency bands.

[0155] For each Mel frequency band, a triangular filter is set up. The frequency response H of the triangular filter for the q-th Mel frequency band at frequency f is... q (f), as follows:

[0156]

[0157] in, and The center Mel frequency is respectively the (q-1)th. The qth center Mel frequency and the (q+1)th center Mel frequency The center frequencies of the (q-1)th, qth, and (q+1)th are obtained by inverse derivation of the Mel transform equation;

[0158] Each spectral sequence in the spectral set F is multiplied sequentially by Q triangular filters and the logarithm is taken to obtain the logarithmic domain signal set LF = {lf1, ..., lf2}. n ,…,lf N}, lf n ={lfn (1),…,lf n (q),…,lf n (Q)} is the spectral sequence f at the nth sampling time. n The logarithmic domain signal sequence obtained by successively multiplying by Q triangular filters and taking the logarithm, where lf n (q) is the spectral sequence f at the nth sampling time. n The logarithmic domain signal obtained by multiplying by the q-th triangular filter and taking the logarithm is given by the following formula: Among them, f n (k) represents the k-th frequency f of the digital audio signal y(n). k The frequency domain signal, H q (f k ) is the triangular filter for the q-th Mel frequency band at frequency f k Frequency response at that location;

[0159] The logarithmic domain signal set LF = {lf1, ...,lf2} is transformed by discrete cosine transform. n ,…,lf N} Transform to the cepstral domain to generate Mel frequency cepstral coefficients MF = [mf1, ..., mf n ,…,mf N ] T ,mf n The Mel-frequency cepstral eigenvector generated by transforming the nth logarithmic domain signal to the cepstral domain is as follows:

[0160] mf n =[mf n (1),…,mf n (u),…,mf n (U)],

[0161] Among them, mf n (u) is the nth Mel frequency cepstral eigenvector mf n The specific calculation formula for the u-th Mel frequency cepstral feature is as follows:

[0162]

[0163] Where u = 1, ..., U, and U is the dimension of the Mel frequency cepstral eigenvector.

[0164] Specifically, the first and second support vector machines need to be pre-trained. A large number of baby cries, human speech, and other types of sounds are collected in advance and converted into Mel-frequency cepstral coefficients (MFs) and labeled as class 1, class 2, and class 3, respectively. When training the first support vector machine, the goal is to distinguish class 1 and the first residue class in high-dimensional space as much as possible to achieve the judgment of baby crying. When training the second support vector machine, all Mel-frequency cepstral coefficients of class 1 are removed, and Mel-frequency cepstral coefficients of class 2 and the second residue class are distinguished in high-dimensional space as much as possible to achieve the judgment of human speech under the premise that it is clear that it is not baby crying. Compared with the synchronous judgment of the two support vector machines, the progressive judgment can reduce the training overhead and accelerate the convergence of the model. The first residue class is the union of class 2 and class 3, class 1 and the first residue class are mutually exclusive, the second residue class is class 3, and class 2 and the second residue class are mutually exclusive.

[0165] Specifically, deep hidden Markov models include deep networks and hidden Markov models;

[0166] The input of the deep network is the Mel frequency cepstral coefficients (MF). The Mel frequency cepstral coefficients (MF) are linearly modulated by multiple layers of linear weights, and the dimension is changed to the number of states of the Hidden Markov Model. The complex relationship between the Mel frequency cepstral eigenvectors in the Mel frequency cepstral coefficients (MF) is learned through the Softmax function, and the posterior probability distribution of the state vector is generated. The posterior probability distribution records the posterior probability of each state of the Hidden Markov Model.

[0167] Hidden Markov Models (HMMs) define state vectors based on phonemes. These state vectors include silence, vocalization initiation, stable vocalization, and vocalization termination states. A state transition probability matrix (STM) describes the probability of transitioning from one state to another. This STM is a square matrix, and its dimension is directly related to the number of states in the state vector. An observation probability distribution represents the probability of observing the Mel-frequency cepstral coefficient (MF) in a single state; this is typically replaced by the posterior probability distribution output by the deep network. An initial state probability distribution is defined as the distribution of the probability of being in each state at the initial time step. At each time step, the probability of transitioning from the current state to the next state and the probability of observing the MF are determined based on the state transition probability matrix and the observation probability distribution. The Viterbi algorithm is used to calculate the maximum probability path from the current state to the remaining states through dynamic programming and records the path information. The optimal state sequence is obtained by backtracking. A pre-determined pronunciation dictionary is consulted, and the speech content z is generated based on the phoneme combination in the optimal state sequence number. The Viterbi algorithm is an existing algorithm.

[0168] Furthermore, the speech content z is split using a context-constrained word segmentation model and reassembled into a speech phrase sequence Ph using a translational sliding window. A skip-word model is then used to convert all speech phrases—including trigger phrases, all functional phrases, and all speech phrases in the speech phrase sequence Ph—into vector representations, and similarity is used for matching and verification. The specific steps include:

[0169] The speech content z is input into the context-constrained word segmentation model, which includes a bidirectional long short-term memory network layer and a conditional random field.

[0170] The bidirectional long short-term memory network layer includes a forward LSTM and a backward LSTM. The forward LSTM processes the speech content z from the beginning to the end, and learns the influence of the previous Chinese character on the next Chinese character through the input gate and the forget gate, capturing the forward semantic association of the speech content z. The backward LSTM processes the speech content z from the end to the beginning, capturing the reverse semantic association of the speech content z. The bidirectional long short-term memory network layer can capture the semantic information of the preceding and following Chinese characters at the same time, and generate semantic features that help to understand the contextual dependencies of the speech content z.

[0171] Conditional Random Fields (CRF) calculate the probability that the feature value of each dimension of semantic features belongs to a single part-of-speech tag based on the state feature function to obtain all possible candidate word segmentation sequences. It then uses a transition feature function to constrain the candidate word segmentation sequences by focusing on the transition relationship between part-of-speech tags, calculates the conditional probability of each candidate word segmentation sequence, and uses the Viterbi algorithm for dynamic programming to find the optimal output order of part-of-speech tags. The speech content z is then labeled sequentially for word segmentation to obtain the word segmentation sequence Wo. Here, CRF is an existing algorithm, and part-of-speech tags include nouns, verbs, adjectives, and adverbs. The transition feature function essentially represents the constraint relationship between part-of-speech tags. For example, the sentence structure "noun + verb + adjective + noun" is generally reasonable, while the sentence structure "noun + noun + adverb" is clearly less reasonable.

[0172] The sliding window width is set according to the maximum phrase length among the trigger phrase and all functional phrases. The sliding window width refers to the number of word segments that can be accommodated between the start edge and the end edge of the sliding window. The start edge is placed in the first word segment in the word segmentation sequence Wo. The word segments within the sliding window width are combined into the first speech phrase. The start edge is moved to the second word segment and combined to generate the second speech phrase. This continues until the end edge reaches the last word segment in the word segmentation sequence Wo, at which point the last speech phrase is generated. All speech phrases are then used to generate the speech phrase sequence Ph in order.

[0173] The skip-word model is used to transform input phrases into vector representations. It extracts the central word from the input phrase and determines the co-occurrence probability of the central word with context words. Based on a pre-defined corpus, the central word is converted into a one-hot encoding. This one-hot encoding is then expanded and reduced in dimensionality using two linear layers, making the dimension equal to the number of words in the input phrase. A Softmax function is used to generate the probability distribution of the input phrase. Backpropagation is employed to adjust the weights of the linear layers to minimize the error between the probability distribution and the co-occurrence probability of the input phrase. Finally, the weights of the first linear layer are used as the vector representation of the input phrase. Specifically, trigger phrases are transformed into trigger vector representations using the skip-word model, each functional phrase is transformed into a corresponding functional vector representation, and each speech phrase in the speech phrase sequence Ph is transformed into a corresponding speech vector representation.

[0174] Calculate the trigger cosine similarity between the trigger vector representation and each speech vector representation, and determine whether there is a trigger cosine similarity greater than or equal to the similarity threshold.

[0175] If it does not exist, it indicates that there is no trigger phrase in the speech phrase sequence Ph, the trigger phrase fails to match the speech phrase sequence Ph, and the sound digital signal sequence y is ignored. Here, the trigger phrase is a preset specific phrase used to determine the validity of the speech phrase sequence Ph and prevent the control system from being miscontrolled due to daily conversation.

[0176] If it exists, it indicates that there is a trigger phrase in the speech phrase sequence Ph. The trigger phrase matches the speech phrase sequence Ph. Further, the function cosine similarity between each function vector representation and each speech vector representation is calculated to determine whether there is a function cosine similarity greater than or equal to the similarity threshold.

[0177] If it does not exist, it means that there are no functional phrases in the speech phrase sequence Ph, all functional phrases fail to match the speech phrase sequence Ph, and the sound digital signal sequence y is ignored;

[0178] If they exist, all functional phrases with a functional cosine similarity greater than or equal to the similarity threshold are considered to match the speech phrase sequence Ph. The instruction encoding table is then searched to obtain the control instructions corresponding to all matched functional phrases.

[0179] Furthermore, the self-control module includes a bedwetting warning unit and a video monitoring unit;

[0180] The bedwetting indicator unit receives a high-level signal y h The system records the reception time and immediately generates a prompt command, which is then sent to the execution module. Starting from the reception time, after each prompt interval, it checks whether the system is still receiving a high-level signal y. h If it is still receiving a high-level signal y hThe prompt command will be repeatedly generated and sent to the execution module. If a high-level signal y is not received... h Then clear the recorded reception time;

[0181] After the control system is activated, the video monitoring unit periodically generates nodes according to the monitoring interval, and extracts a video digital signal sequence Y at each node. vid All video digital signals from the previous node to the current node are converted into time-segment video using the H.265 encoding algorithm. Video analysis of time periods using facial recognition and optical flow detection The system analyzes the infant's facial and physical features to identify the infant's sleep state. Based on the rule engine, it generates control instructions or prompts and transmits them to the execution module. The infant's sleep state includes deep sleep, light sleep, semi-awake, awake, and abnormal sleep.

[0182] like Figure 3 As shown, further analysis of time-segment videos using facial recognition and optical flow detection is conducted. Identifying an infant's sleep state involves the following specific steps:

[0183] Video of a specific time period Equal-interval sampling is performed to acquire a total of T sampled frame images to construct a sampled frame image sequence. in, Let be the sampled frame image of frame t;

[0184] Combined with the Dlib library for sampling frame image sequences Eye state recognition is performed. The baby's face region in each sampled frame image is located by the face detector of Dlib. The coordinates of key points around the eyes in each sampled frame image are obtained by using the 68 facial key point models pre-trained by Dlib. The face detector and the 68 facial key point models are existing models.

[0185] Calculate the vertical distance between the midpoint of the upper eyelid and the midpoint of the lower eyelid in each sampled frame image and take the average value to obtain the sampled frame image sequence. The average vertical distance is used as the ratio of the average vertical distance to the maximum vertical distance. The maximum vertical distance is determined by averaging the vertical distances of a large number of infants when their eyes are fully open. The eye-opening ratio is always within the range of 0 to 1.

[0186] Building upon the lightweight MobileNetV3 model, a spatial attention mechanism is introduced to improve the accuracy of facial expression recognition. The initial convolutional layer in the MobileNetV3 model extracts the sampled frame image sequence. For each sampled frame image, the spatial attention mechanism performs average pooling and global pooling on the feature map of the RGB three-channel sub-image, concatenates the two pooling results, and learns the attention weights of different spatial locations in the sampled frame image through convolution operation, and obtains the spatial attention map through the Sigmoid function.

[0187] The spatial attention map is connected to the depthwise separable convolutional layer of the MobileNetV3 model. The spatial attention map is upgraded, features are extracted and reduced in dimensionality through pointwise convolution, depthwise convolution and pointwise convolution in sequence. It is then transformed into a spatial attention vector through global average pooling. After linear modulation through a fully connected layer, the expression probability distribution of the sampled frame images is generated by the Softmax function. The average expression probability distribution of all sampled frame images is calculated. The expression level includes calm, moderate and rich.

[0188] Detection of sampled frame image sequences using optical flow method The temporal changes of pixels are defined by a luminance identity, which describes the frame image of frame t. The pixel with pixel coordinates (α1, α2) in the image. The brightness Lg(α1,α2,t) undergoes a pixel coordinate shift from frame t to frame t+1, with an shift amount of (Δα1,Δα2). This shift occurs in the frame image of frame t+1. The corresponding pixel becomes The brightness Lg(α1+Δα1,α2+Δα2,t+1) in the (t+1)th frame is still the same as the brightness Lg(α1,α2,t), i.e., Lg(α1,α2,t)=Lg(α1+Δα1,α2+Δα2,t+1). Furthermore, based on the assumption of spatiotemporal continuity, the offset (Δα1,Δα2) will not be too large. The image sequence describing the sampled frames is found through TV-L1 iterative optimization. Optical flow feature sequence, optical flow feature map in optical flow feature sequence and sampled frame image sequence The sampled frame images correspond one-to-one. The optical flow feature map of frame t is a two-dimensional vector field that describes the frame images from frame (t-1). Frame image up to frame t The TV-L1 algorithm iteratively optimizes the movement trajectory and speed of pixels over time into the existing algorithm.

[0189] Optical flow feature sequences are aggregated through temporal convolutional layers. The temporal convolutional layers capture the temporal changes of optical flow features in the optical flow feature sequence through two identical residual blocks. Each residual block consists of dilated convolution and residual structure. The dilated convolution inserts zero elements between the convolution kernels to expand the receptive field. The residual structure superimposes the output of the dilated convolution with the input. After passing through two identical residual blocks, the fully connected layer reduces the dimension and outputs an aggregated optical flow feature map with the same dimension as any optical flow feature map in the optical flow feature sequence.

[0190] The aggregated optical flow feature map is linearly reduced in dimensionality by a fully connected layer, and the motion amplitude probability distribution is generated again by the Softmax function, which includes small, medium and large motion amplitudes.

[0191] Set up a scoring system, defining calm, moderate, and rich facial expressions as scores of 0, 0.5, and 1 respectively, and small, medium, and large movement amplitudes as scores of 0, 0.5, and 1 respectively. The facial expression scores are weighted and summed according to the probability distribution of facial expression levels to obtain the facial expression score, and the movement amplitude scores are weighted and summed according to the probability distribution of movement amplitude to obtain the movement score. Both facial expression scores and movement scores are within the range of 0-1.

[0192] Based on the reference weights of eye opening, facial expression, and movement, the proportion of eye opening, facial expression score, and movement score are weighted and summed again to obtain a comprehensive sleep score. The infant's sleep state is determined based on the comprehensive sleep score interval. Here, the reference weights and comprehensive score intervals are predefined. In this embodiment, the reference weights are 0.3:0.3:0.4, and the specific comprehensive score intervals are as follows:

[0193] A comprehensive score of less than 0.1 corresponds to the infant's sleep state being deep sleep;

[0194] A comprehensive score greater than or equal to 0.1 and less than 0.35 corresponds to a light sleep state in infants.

[0195] A comprehensive score greater than or equal to 0.35 and less than 0.65 corresponds to an infant's sleep state of semi-awakeness;

[0196] A comprehensive score greater than or equal to 0.65 and less than 0.9 corresponds to an infant's awake sleep state;

[0197] A comprehensive score of 0.9 or higher indicates an abnormal sleep state in infants.

[0198] Specifically, the rule engine generates corresponding control instructions or prompts based on the infant's sleep state, including:

[0199] If the baby is in deep sleep, a series of control commands are generated and sent to the execution module according to the current working status of the variable speed synchronous motor and the speaker, so that the variable speed synchronous motor stops working and the speaker stops playing music and stories. For example, if the variable speed synchronous motor is currently running and the speaker is only playing music, only the swing stop command and the music stop command need to be generated.

[0200] If the baby is in a light sleep state, control commands are generated sequentially based on the current working status of the variable speed synchronous motor and the speaker to make the variable speed synchronous motor work at a low speed and the speaker stop playing music and stories. For example, if the variable speed synchronous motor is not currently started and the previous motor gear was high, and the speaker is currently started and playing music and stories, then the swing start command, two swing deceleration commands, music stop command and story stop command are generated sequentially.

[0201] If the baby is in a semi-awake state, control commands are generated sequentially based on the current working status of the variable speed synchronous motor and the speaker to make the variable speed synchronous motor work at the medium speed and the speaker only play music. For example, if the variable speed synchronous motor has started and the current motor speed is exactly the medium speed, and the speaker is playing music and a story at the same time, then only the story stop command needs to be generated.

[0202] If the baby is awake, control commands are generated sequentially based on the current operating status of the variable speed synchronous motor and the speaker to make the variable speed synchronous motor work at a high speed and the speaker play music and stories at the same time. For example, if the variable speed synchronous motor has started and the current motor gear is at a low speed, and the speaker has not started, then two swing acceleration commands, music playback commands, and story playback commands are generated sequentially.

[0203] If the baby's sleep state is abnormal, a prompt instruction will be generated and sent to the execution module.

[0204] Furthermore, the execution module includes a power switching unit, a variable speed synchronous motor execution unit, a sound playback unit, and a cloud communication unit;

[0205] The power switching unit receives a power-on command or a power-off command, and turns the circuit switch on or off to control the start and stop of the control system and the controlled device.

[0206] The variable speed synchronous motor execution unit receives the swing start command, starts the variable speed synchronous motor according to the previous motor gear, receives the swing stop command, records the current motor gear and shuts down the variable speed synchronous motor, receives the swing acceleration command, determines whether the current motor gear is high. If so, it ignores it; otherwise, it raises the motor gear by one gear. It receives the swing deceleration command, determines whether the current motor gear is low. If so, it ignores it; otherwise, it lowers the motor gear by one gear. The motor gears include low, medium and high.

[0207] The sound playback unit receives music playback instructions or story playback instructions, sequentially calls the next built-in music or story based on the previously played music or story, and plays it through the speaker. Upon receiving a music stop instruction or story stop instruction, it records the currently playing music or story and turns off the speaker.

[0208] The cloud communication unit receives the video digital signal sequence Y. vid The video database is simultaneously uploaded to the cloud platform. The system receives prompts and sends them to all paired devices through the cloud platform. The cloud platform allocates different storage areas and establishes corresponding video databases based on the controlled device numbers controlled by the control system. Only paired devices can view the video database of the corresponding control system through the cloud platform.

[0209] Example 2

[0210] like Figure 4 As shown, the present invention also discloses a controlled device controlled by the control system, including a cradle body 1, a base 2, a drive device 3, a universal ball bearing 4, a universal rocking plate 5, a cradle body connecting flange 6, a limiting rubber belt 7, a urine sensor 8, and an integrated processing device 9.

[0211] Cradle 1 is used to accommodate a baby while sleeping;

[0212] The base 2 is used to increase the stability of the controlled device, and there is a compartment inside the base 2;

[0213] The drive unit 3 is installed in the compartment of the base 2 and includes a variable speed synchronous motor 31 and a variable speed synchronous motor rocker arm 32. When the variable speed synchronous motor 31 starts, it drives the variable speed synchronous motor rocker arm 32 to swing in the left and right and forward and backward directions. The variable speed synchronous motor rocker arm 32 is fixedly connected to the universal rocker plate 5.

[0214] The universal ball bearing 4 is used to support the universal rocking disk 5, reduce the friction between the universal rocking disk 5 and the base 2, and make the swing of the universal rocking disk 5 smoother.

[0215] The universal swing disk 5 and the variable speed synchronous motor rocker arm 32 swing in the same direction.

[0216] The cradle body connecting flange 6 connects the universal rocking plate 5 and the cradle body 1, so that the cradle body 1 and the universal rocking plate 5 keep swinging in the same direction.

[0217] The limiting rubber belt 7 connects the base 2 and the universal swing plate 5 to prevent the universal swing plate 5 from derailing after being hit by external force.

[0218] Urine sensor 8 is installed at the bottom of cradle body 1 to sense urine contact and generate a high-level signal y. h ;

[0219] The integrated processing unit 9 is used to acquire audio digital signal sequences y and continuously monitor and acquire video digital signal sequences y. vid It provides users with four control methods: touch control, remote control, voice control, and automatic control, to control the start, stop, and gear adjustment of the drive device 3, as well as the playback and stop of music and stories. Simultaneously, it uploads the video digital signal sequence Y to the cloud platform. vid And prompts / instructions.

[0220] Furthermore, the integrated processing device 9 includes a keypad screen, a rear camera, a microphone array, a speaker, a wireless communication module, and a processor;

[0221] The keypad screen is located on the front of the integrated processing unit 9, displaying the keys and providing touch control for the user;

[0222] The rear camera is fixed to the back of the integrated processing unit 9 to monitor and record the inside of the cradle body 1 from a specific angle;

[0223] Microphone arrays are used to collect analog sound signals x raw And through filtering and analog-to-digital conversion, a digital sound signal sequence y is generated;

[0224] The speaker is controlled by the processor to play or stop playing music or a story.

[0225] The wireless communication module receives remote control signals and forwards them to the processor, and receives the video digital signal sequence Y. vid Provide prompts and instructions and upload them to the cloud platform;

[0226] The processor at the high level signal y h During reception, prompts are continuously generated; user touch buttons are identified to generate control commands; remote control signals are received to generate control commands; an audio digital signal sequence y is received and its characteristics (e.g., infant crying or human speech) are determined; if it is infant crying, a prompt command is generated; if it is human speech, it is converted into speech content z and matched with trigger phrases and all function phrases. Upon successful trigger phrase matching, the corresponding control command is generated by retrieving the command encoding table based on the matched function phrase. The video digital signal sequence Y is also received. vid It also periodically extracts and transcodes video into time segments. The system identifies the baby's sleep state through facial recognition and optical flow detection, and generates control commands or prompts based on the rule engine. The control commands are used to control the start, stop and gear adjustment of the drive device 3, as well as the playback and stop of music and stories from the speaker. The prompts are then handed over to the wireless communication module.

[0227] This invention discloses a control system and a controlled device, including a data acquisition module, a controlled module, a self-control module, and an execution module. The data acquisition module continuously sends a high-level signal in the presence of urine, acquires analog sound signals, processes them to generate a digital sound signal sequence, and continuously acquires a digital video signal sequence. The controlled module senses user touch or receives remote control signals to generate control commands, converts the digital sound signal sequence into Mel-frequency cepstral coefficients, and identifies whether the digital sound signal sequence is an infant's cry or human speech based on a progressive binary classification model. If it is an infant's cry, a prompt command is generated; if it is human speech, a deep hidden Markov model combined with a context-constrained word segmentation model is used to convert the Mel-frequency cepstral coefficients into speech content and split and reassemble it into a speech phrase sequence based on semantic dependencies, utilizing a skip-word model. The system converts the voice phrases in the trigger phrase, function phrase, and voice phrase sequences into vector representations and performs matching verification. When the trigger phrase matches successfully and at least one function phrase matches successfully, control commands are generated based on the matched function phrases; otherwise, the audio digital signal sequence is ignored. The automatic control module periodically captures video digital signal sequences and transcodes them into time-segment videos. It identifies the baby's sleep state through facial recognition and optical flow detection, generates control commands or prompts based on the rule engine, and periodically generates prompts during high-level signal reception. The execution module executes the control commands, receives the video digital signal sequences and prompts, and simultaneously uploads them to the cloud platform. This achieves a four-in-one collaborative control method integrating touch, remote, voice, and automation, improving the intelligence level of the baby cradle.

[0228] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A control system, characterized in that, Includes controlled modules and automatic control modules; The controlled module senses user touch or receives remote control signals to generate control commands, receives audio digital signal sequences and converts them into Mel-frequency cepstral coefficients, identifies whether the audio digital signal sequence is an infant's cry or human language based on a progressive binary classification model, generates prompt commands if it is an infant's cry, and generates prompt commands if it is human language if it is human language. It uses a deep hidden Markov model and a context-constrained word segmentation model to analyze Mel-frequency cepstral coefficients to generate speech content and splits and reassembles it into speech phrase sequences based on semantic dependencies. It uses a skip-word model to match and verify the speech phrase sequences to detect whether they match trigger phrases and functional phrases. When a trigger phrase and at least one functional phrase are matched, control commands are generated based on the matched functional phrases. Otherwise, the audio digital signal sequence is ignored. The self-control module periodically captures video digital signal sequences and identifies the baby's sleep state through transcoding, face recognition, and optical flow detection. It generates control commands or prompts based on the rule engine and continuously generates prompts during high-level signal reception. The step of using a skip-word model to match and verify speech phrase sequences to detect whether they match trigger phrases and functional phrases includes the following specific steps: The method utilizes a skip-word model to extract the central word of the input phrase and determine the co-occurrence probability of the central word and the context words. Based on the corpus, the central word is converted into one-hot encoding. The probability distribution of the input phrase is generated through two linear layers and a Softmax function. Backpropagation is used to adjust the weights of the linear layers to minimize the error between the probability distribution and the co-occurrence probability of the input phrase. The weights of the first linear layer are used as the vector representation of the input phrase. The input phrase includes trigger phrases, function phrases, and speech phrases in the speech phrase sequence. The vector representation includes trigger vector representation, function vector representation, and speech vector representation. Calculate the trigger cosine similarity between the trigger vector representation and each speech vector representation, and simultaneously calculate the functional cosine similarity between each function vector representation and each speech vector representation, and compare them sequentially with the similarity threshold. If both trigger cosine similarity and function cosine similarity are greater than or equal to the similarity threshold, corresponding control instructions are generated based on all function phrases whose function cosine similarity is greater than or equal to the similarity threshold. If neither trigger cosine similarity nor function cosine similarity is greater than or equal to the similarity threshold, the audio digital signal sequence is ignored.

2. The control system as described in claim 1, characterized in that, The controlled module includes a voice control unit; The voice control unit filters and logarithmically transforms the audio digital signal sequence in the frequency domain with a Mel filter bank to generate Mel frequency cepstral coefficients. A progressive binary classification model is used to sequentially determine whether the audio digital signal sequence represents an infant's cry or human speech based on the Mel frequency cepstral coefficients. If it represents an infant's cry, a prompt command is generated; if it represents human speech, a deep hidden Markov model is used to convert the Mel frequency cepstral coefficients into speech content. A context-constrained word segmentation model is used to split the speech content into word sequences based on semantic dependencies and constructs speech phrase sequences by combining them through a translational sliding window. A skip-word model is used to match and verify trigger phrases, functional phrases, and speech phrase sequences. If a trigger phrase and at least one functional phrase match the speech phrase sequence, a corresponding control command is generated based on the matched functional phrase; otherwise, the audio digital signal sequence is ignored.

3. The control system as described in claim 2, characterized in that, The process of filtering and performing logarithmic operations on the audio digital signal sequence in the frequency domain with a Mel filter bank to convert it to the cepstral domain and generate Mel frequency cepstral coefficients includes the following specific steps: The audio digital signal sequence is processed by Discrete Fourier Transform to generate a spectrum set, which includes... The spectral sequence at the nth sampling time, the th The spectral sequence at each sampling time includes A frequency domain signal of a certain frequency. and These represent the total number of sampling times and the total number of frequencies, respectively. Based on the Mel transform equation, the frequency range of the audio digital signal sequence is transformed into the Mel frequency range and divided into... Each Mel frequency band is determined separately. The center Mel frequency of the Mel frequency band The total number of Mel frequency segments; Based on the center Mel frequency of the three adjacent Mel frequency bands, a triangular filter is set in the middle Mel frequency band, and a total of [number] Mel frequency bands are obtained. Each triangular filter corresponds one-to-one with a Mel frequency band. In the spectrum set The spectral sequences at each sampling time are respectively compared with... Multiplying the triangular filters and taking their logarithms generates a logarithmic domain signal set, which includes... The logarithmic domain signal sequence at the nth sampling time, the i-th The logarithmic domain signal sequence at each sampling time includes The logarithmic domain signals, where the first... The logarithmic domain signal is the first Substituting each frequency in the spectral sequence at the sampling time into the... The sum of the product of the frequency response of each triangular filter and the corresponding frequency signal. ; The logarithmic domain signal set is transformed to the cepstral domain using discrete cosine transform, generating Mel-frequency cepstral coefficients, which include... Mel frequency cepstral eigenvectors at each sampling time.

4. The control system as described in claim 2, characterized in that, The deep hidden Markov model includes deep networks and hidden Markov models; The deep network linearly modulates the Mel frequency cepstral coefficients through linear layers and the Softmax function and maps them to generate the posterior probability distribution of the state vector. The posterior probability distribution records the posterior probability of each state in the state vector. The Hidden Markov Model (HMM) defines a state vector based on phonemes. The state vector includes a silent state, a vocalization start state, a stable vocalization state, and a vocalization end state. A state transition probability matrix is ​​defined to describe the probability of transitioning from one state to another. The posterior probability distribution of the state vector reflects the probability of observing Mel-frequency cepstral coefficients in a single state. An initial state probability distribution is defined for the initial time step. At each time step, the probability of transitioning from the current state to the next state is determined based on the state transition probability matrix, and the probability of observing Mel-frequency cepstral coefficients is determined based on the posterior probability distribution of the state vector. The Viterbi algorithm is used to calculate the maximum probability path from the current state to the remaining states and the path information is recorded. The optimal state sequence is obtained by backtracking, and the phonemes in the optimal state sequence are combined according to the pronunciation dictionary to generate speech content.

5. The control system as described in claim 2, characterized in that, The context-constrained word segmentation model includes a bidirectional long short-term memory network layer and a conditional random field. The bidirectional long short-term memory network layers employ forward LSTM and backward LSTM respectively. Through input gate and forget gate, the influence between adjacent Chinese characters in the speech content is learned in two orders to capture the semantic associations in forward and reverse order and generate semantic features. Conditional random fields (CRFs) calculate the probability of each feature value in semantic features belonging to a single part-of-speech tag based on the state feature function to obtain all candidate word segmentation sequences. By focusing on the transition relationship between part-of-speech tags through the transition feature function, the candidate word segmentation sequences are constrained. The conditional probability of each candidate word segmentation sequence is calculated. The optimal part-of-speech tag output order is found through the Viterbi algorithm, and the speech content is labeled for word segmentation to generate word segmentation sequences.

6. The control system as described in claim 1, characterized in that, The self-control module includes a bedwetting warning unit and a video monitoring unit; The bedwetting warning unit receives a high-level signal and records the reception time, periodically generates a warning instruction, and clears the reception time when no high-level signal is received. After the intelligent control system is turned on, the video monitoring unit periodically captures video digital signal sequences and converts them into time-segment videos according to the H.265 encoding algorithm. It identifies the baby's sleep state by analyzing the time-segment videos through facial recognition and optical flow detection. Based on the correspondence between the baby's sleep state and the target working state in the rule engine, it generates control instructions or prompts in combination with the current working state to adjust the current working state to the target working state. The baby's sleep state includes deep sleep, light sleep, semi-awake, awake, and abnormal.

7. The control system as described in claim 6, characterized in that, The method of identifying infant sleep status through facial recognition and optical flow detection analysis of video over a period of time includes the following specific steps: The video during the sampling period is used to construct a sequence of sampled frame images. The coordinates of the infant's face region and key points around the eyes in each sampled frame image are located by a face detector and a model of 68 facial key points. The vertical distance between the midpoint of the upper eyelid and the midpoint of the lower eyelid is calculated. The average vertical distance is obtained and divided by the maximum vertical distance to generate the proportion of eyes open. The MobileNetV3 model extracts feature maps from each sampled frame image in the sampled frame image sequence through convolutional layers, and uses a spatial attention mechanism to concatenate the average pooling results and global pooling results of the feature maps in multiple channels. It then generates the corresponding spatial attention map through convolution and the Sigmoid function. The MobileNetV3 model performs dimensionality upscaling, feature extraction, and dimensionality reduction on the spatial attention map of each sampled frame image. It then generates the corresponding probability distribution of facial expression level through global average pooling, fully connected layers, and the Softmax function, and calculates the average probability distribution of facial expression level, which includes calm, moderate, and rich facial expression levels. Optical flow is used to detect the temporal changes of pixels in a sampled frame image sequence. A brightness identity is defined to describe the constant brightness of corresponding pixels between adjacent frames. Based on the assumption of spatiotemporal continuity, the optical flow feature sequence describing the sampled frame image sequence is found through TV-L1 iterative optimization. Optical flow feature sequences are input into a temporal convolutional layer. The temporal convolutional layer captures the temporal changes of optical flow features through two identical residual blocks. Dimensionality reduction and mapping are performed through multiple fully connected layers and a Softmax function to generate a probability distribution of action amplitude. The residual block includes dilated convolution and residual structure. Dilated convolution inserts zero elements between convolution kernels to expand the receptive field. The residual structure superimposes the output of the dilated convolution with the input. Action amplitudes include small, medium, and large. Based on the scores for each level of facial expression intensity and movement amplitude, the facial expression intensity probability distribution and movement amplitude probability distribution are weighted and summed to generate facial expression score and movement score. The eye opening ratio, facial expression score and movement score are then weighted and summed again according to reference weights to generate a comprehensive sleep score. The infant's sleep state is determined based on the comprehensive sleep score range.

8. The control system as described in claim 1, characterized in that, The intelligent control system also includes a data acquisition module and an execution module; The acquisition module includes a urine sensing unit, a sound sensing unit, and a video monitoring unit. The urine sensing unit senses urine contact based on the principle of capacitive sensing and the conduction characteristics of an NMOS transistor. It generates a high-level signal when urine comes into contact and transmits it to the automatic control module. It stops generating the high-level signal when the urine is cleared. The sound sensing unit collects analog sound signals, which are then converted into a digital sound signal sequence by passing them through a filter bank, an analog-to-digital converter, and an adaptive filtering algorithm, and transmitted to the controlled module. The video monitoring unit collects ambient light intensity in real time and adaptively determines the supplementary lighting brightness. It continuously captures and generates monitoring video, which is converted into a digital video signal sequence using the H.265 encoding algorithm and sent synchronously to the controlled module and the execution module. The execution module executes the opening and closing of circuit switches, the opening and closing of variable speed synchronous motors and gear adjustment, the playback and stopping of music and stories based on control commands, and sends the received video digital signal sequence and prompt commands to the cloud platform. The cloud platform is used to establish remote communication between the intelligent control system and the paired device.

9. A controlled device, controlled by the control system according to any one of claims 1-8, comprising a cradle body for accommodating a sleeping infant, a base for increasing stability, a universal rocking disc for swinging, a cradle body connecting flange for ensuring that the cradle body and the universal rocking disc swing in the same direction, and a limiting rubber band for preventing the universal rocking disc from derailing, characterized in that, The controlled device also includes a drive unit, a universal ball bearing, a urine sensor, and an integrated processing unit. The drive device is controlled by an integrated processing device and is connected to the universal swing disk. When started, it swings left and right and forward and backward to drive the universal swing disk to swing in the same direction. The universal ball bearing is used to support the universal sway disk and reduce the friction between the universal sway disk and the base. The urine sensor is installed at the bottom of the cradle body and is used to sense urine contact and generate a high-level signal; The integrated processing device is used to acquire audio digital signal sequences and continuously monitor and acquire video digital signal sequences. It provides four control methods: user touch control, remote control, voice control, and automatic control. It is used to control the start, stop, and gear adjustment of the drive device, as well as the playback and stop of music and stories. At the same time, it uploads video digital signal sequences and prompts to the cloud platform.

Citation Information

Patent Citations

  • Voice recognition method and device

    CN105702250A

  • Cloud automatic camera monitoring system and method facilitating infant nursing

    CN109982046A