Smart home scene interaction method based on image recognition technology

The method improves smart home interaction by using image recognition to analyze joint positions and track motion sequences, enhancing classification stability and personalizing environmental adjustments through feedback loops.

CN120318907APending Publication Date: 2025-07-15SHENZHEN LIGUAN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510449013.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, motion analysis does not go deep into the joint level, and it is difficult to accurately distinguish similar actions, and it is prone to misclassification or action recognition errors. The equipment control instructions are executed in one-way, and they cannot be effectively optimized in combination with device feedback, resulting in failed command execution or deviation in adjustment effects.

Method used

By capturing image data based on the camera, using the OpenPose algorithm to identify joint positions, combining time series analysis and environmental parameters, generating behavior pattern sequences, performing posture classification, dynamically adjusting the home environment, and forming closed-loop control.

Benefits of technology

It improves the stability and accuracy of behavior classification, can identify subtle movement changes, realize personalized adjustment, and enhances the adaptability of smart home systems to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318907A_ABST
    Figure CN120318907A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a smart home scene interaction method based on an image recognition technology, which comprises the following steps of: capturing real-time image data of a resident based on a camera, and extracting pixel points and color depth information in an image to obtain initial posture data; and based on the initial attitude data, an OpenPose algorithm is applied to detect a joint position, the attitude of the resident is identified, and a joint position analysis result is generated. According to the invention, based on the motion trail analysis of the time sequence, the system can analyze the conversion process of different postures, and the stability of behavior classification is improved. And joint motion analysis is adopted, so that classification of behavior states is more detailed, different subtle motion changes under the same category can be recognized, and misjudgment is reduced. In combination with an environment parameter dynamic adjustment strategy, the system can perform personalized adjustment according to a user state instead of depending on a fixed rule to perform equipment control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a smart home scenario interaction method based on image recognition technology. Background Art

[0002] Image recognition technology is a core branch in the field of computer vision, which involves parsing and understanding the content of images captured from cameras or other image acquisition devices through algorithms and models. The smart home scenario interaction method based on image recognition technology involves applying image recognition technology to the smart home system to improve the interactivity and automation level of the home environment.

[0003] In the prior art, motion analysis does not go deep into the joint level, making it difficult to accurately distinguish similar actions, and prone to misclassification or incorrect action recognition. At the same time, device control instructions are often executed unidirectionally and fail to effectively combine device feedback for dynamic optimization, which may lead to failed instruction execution or deviation in adjustment effects. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies in the prior art and propose a smart home scenario interaction method based on image recognition technology.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions. The smart home scenario interaction method based on image recognition technology includes the following steps:

[0006] Based on the real-time image data of the occupant captured by a camera, extract the pixel points and color depth information in the image to obtain preliminary pose data; based on the preliminary pose data, apply the OpenPose algorithm to detect the joint positions, identify the pose of the occupant, and generate a joint position analysis result.

[0007] Based on the joint position analysis result, perform serialization processing on the motion trajectory of each joint through time series analysis to obtain a behavior pattern sequence; based on the behavior pattern sequence, analyze the behavior state of the occupant, where the behavior state includes standing, sitting / lying, or falling, and generate a pose classification result.

[0008] Based on the pose classification result, decide whether it is necessary to adjust the home environment, including adjusting the lights or air conditioner, and generate a preset trigger decision.

[0009] Based on the preset trigger decision, send adjustment commands to the smart home devices through a controller, including setting the indoor light brightness, air conditioner temperature, and curtain state, to obtain the smart home response status.

[0010] Preferably, the step of obtaining the preliminary pose data is as follows:

[0011] Capture real-time image data of the occupant through a camera, perform color and brightness correction on each frame of the image, and obtain corrected real-time video data;

[0012] Based on the corrected real-time video data, apply bicubic interpolation to enhance the resolution of each frame of the image to obtain preliminary pose data.

[0013] Preferably, the steps for obtaining the joint position analysis result are as follows:

[0014] Based on the preliminary pose data, use the OpenPose algorithm to detect the joint feature points in the image to obtain joint detection data;

[0015] According to the joint detection data, obtain the spatial coordinates of each feature point, match the known joint connection patterns, perform connectivity evaluation on all estimated joint combinations, and eliminate abnormal connection structures to generate joint position data;

[0016] Based on the joint position data, analyze the spatial distribution state of each joint, judge the angular relationship between adjacent joints, match the established human pose patterns, and generate joint position analysis results.

[0017] Preferably, the steps for obtaining the behavior pattern sequence are as follows:

[0018] Based on the joint position analysis result, extract the time series data of all joints, arrange the spatial coordinates of each joint in chronological order, smooth the joint movements in consecutive frames, and generate joint movement trajectory data;

[0019] According to the joint movement trajectory data, calculate the curvature change rate of the joint trajectory. The calculation formula is:

[0020]

[0021] where S is the curvature change rate, Q j is the position coordinate of the joint at time T j moment, and k is the total number of frames in the time series;

[0022] Based on the curvature change rate, analyze the time series movement patterns of all joints, classify according to the change trend of the trajectory, and form a behavior pattern sequence.

[0023] Preferably, the steps for obtaining the pose classification result are as follows:

[0024] Based on the behavior pattern sequence, extract the joint movement states within all time segments, traverse all time windows to calculate the relative displacement vectors of the joints, and calculate the relative rotation matrix of the joint points based on the three-dimensional Euclidean space transformation to determine the angle change trend between each time point, and generate behavior feature data;

[0025] Calculate the posture classification score according to the behavior characteristic data, and the calculation formula is:

[0026]

[0027] Among them, F is the posture classification score, A s is the joint angle change rate at time s, B s is the joint acceleration change value at time s, C s is the Euclidean distance between adjacent joints at time s, D s is the trajectory curvature at time s, E s is the inertia change rate at time s, and p is the total number of frames within the time window;

[0028] Based on the posture classification score, compare with the known posture model database, analyze the classification characteristics of different posture states in combination with the motion mode set, match the states of standing, sitting / lying, or falling, and generate the posture classification result.

[0029] Preferably, the obtaining step of the preset trigger decision is:

[0030] Based on the posture classification result, extract the posture state data within all time periods, analyze the distribution of different postures in the time series, calculate the duration of each posture, and segment and mark the continuously changing states. Extract the posture anomaly data in combination with the indoor environment parameters to generate the environmental adjustment requirement data;

[0031] According to the environmental adjustment requirement data, calculate the environmental adjustment decision score, and the calculation formula is:

[0032]

[0033] Among them, V is the environmental adjustment decision score, P c is the duration of the current posture, P d is the duration of the previous posture, CQ x is the current light intensity, CQ y is the light intensity recorded last time, R is the change rate of the indoor temperature, S is the change rate of the humidity, and CT is the preset response time interval;

[0034] Based on the environmental adjustment decision score, decide whether to perform the adjustment of the light brightness, air conditioner temperature, or curtain state, and generate the preset trigger decision.

[0035] Preferably, the obtaining step of the smart home response state is:

[0036] Based on the preset trigger decision, parse the home environment adjustment command, extract the target device, adjustment parameters and execution timing in the control instruction, convert the command into an instruction format recognizable by the device, and generate controller instruction data;

[0037] According to the controller instruction data, send instructions to the home intelligent devices, execute the adjustment of the light brightness, air conditioner temperature and curtain state, monitor the execution feedback of the devices, verify the correct transmission and execution of the instructions, and generate device response data.

[0038] Preferably, the step of obtaining the smart home response status further includes: based on the device response data, analyze the operating status of the home intelligent devices, determine whether the adjustment meets the preset requirements, record the status data of the devices after adjustment, and resend instructions to the devices that have not been executed or have abnormal responses to form the smart home response status.

[0039] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0040] In the present invention, based on the motion trajectory analysis of the time series, the transformation process of different postures can be parsed, and the stability of behavior classification can be improved. By using joint motion analysis, the classification of behavior states is more refined, and different subtle motion changes under the same category can be recognized, reducing misjudgment. Combining the dynamic adjustment strategy of environmental parameters, personalized adjustment can be made according to the user's state, rather than relying on fixed rules to control the devices. The control instruction and the device feedback mechanism form a closed loop, and adaptive adjustment can be made according to the execution situation of the home devices, improving the adaptability of the smart home system to complex scenarios. The environmental adjustment strategy comprehensively judges according to the posture classification result and the persistence of the behavior pattern, enabling the home adjustment decision to consider short-term trends and long-term habits, and enhancing the accuracy of home automation. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0043] Please refer to Figure 1 , the present invention provides a technical solution, a smart home scenario interaction method based on image recognition technology, including the following steps:

[0044] Based on the camera capturing the real-time image data of the occupant, extract the pixel points and color depth information in the image to obtain preliminary pose data; based on the preliminary pose data, apply the OpenPose algorithm to detect the joint positions, identify the pose of the occupant, and generate the joint position analysis result;

[0045] Based on the joint position analysis result, serialize the motion trajectory of each joint through time series analysis to obtain the behavior pattern sequence; based on the behavior pattern sequence, analyze the behavior state of the occupant, where the behavior state includes standing, sitting / lying down, or falling, and generate the pose classification result;

[0046] Based on the pose classification result, decide whether it is necessary to adjust the home environment, including lighting or air conditioning adjustment, and generate a preset trigger decision;

[0047] Based on the preset trigger decision, send adjustment commands to the home intelligent devices through the controller, including settings for indoor lighting brightness, air conditioning temperature, and curtain status, to obtain the smart home response status.

[0048] The steps for obtaining the preliminary pose data are as follows:

[0049] Capture the real-time image data of the occupant through the camera, perform color and brightness correction on each frame of the image to obtain the corrected real-time video data;

[0050] Based on the corrected real-time video data, apply bicubic interpolation to enhance the resolution of each frame of the image to obtain the preliminary pose data.

[0051] Specifically, taking the image of the occupant captured in real time by the camera as the initial input, first extract the three benchmark brightness parameters of red, green, and blue from the collected environmental light condition records and denote them as R ref 、G ref 、B ref , and then perform difference calculation for each pixel of each frame of the image to obtain ΔR = |R i -R ref |, ΔG = |G i -G ref |, ΔB = |B i -B ref |, where R i 、G i 、B iThey are the red, green, and blue channel values of the current pixel respectively. If the difference in any of them exceeds 30, it is considered that the brightness deviates too much from the benchmark. Here, 30 is an empirical threshold determined by repeatedly measuring and statistically averaging the deviation of 500 sample images taken under different lighting conditions. Then, the pixels exceeding the threshold are iteratively adjusted according to the normalization correction coefficient, and the brightness difference is distributed to each channel to correct the color and brightness. The initial value of the correction coefficient is set to 1.0 and is gradually fine-tuned in steps of 0.1 according to the difference between the pixel and the reference value. When the difference between the pixel and the benchmark is less than 5 during the iteration, the adjustment stops. This 5 is also the end standard for fine-tuning determined based on actual multiple tests. Finally, the pixel data of each frame after the above processing is recombined to form a continuous image sequence, and the corrected real-time video data is obtained.

[0052] Based on the corrected real-time video data and assuming that interpolation magnification is required in both the horizontal and vertical directions simultaneously, the scaling coefficient α is set to 1.5. According to the principle of bicubic interpolation, for each target pixel to be interpolated, 16 adjacent pixel points in the upper, lower, left, and right rows and columns around it are selected. The weighting factors in the horizontal and vertical directions of each adjacent pixel are calculated, and their product is used as the final interpolation coefficient. The generation process of the weighting factors mainly judges the corresponding relationship between the relative distance and the pixel weight through the interpolation kernel function. For example, when the normalized distance is less than 0.05, no further accumulation is performed and it is directly regarded as a negligible term. The setting of 0.05 here is an empirical range selected after comparing multiple groups of interpolation results. Subsequently, the color values of each adjacent pixel are multiplied by the corresponding interpolation coefficient and accumulated to obtain the new value of the target pixel. For the pixels at the interpolation boundary, if the number of their neighbors is less than 16, the corresponding rows and columns are supplemented within the valid range of the image. In this way, the pixel interpolation magnification operation for each frame is completed repeatedly, and an image sequence for subsequent recognition is generated by continuously processing all frames, obtaining the preliminary pose data.

[0053] The steps to obtain the joint position analysis result are as follows:

[0054] Based on the preliminary pose data, the OpenPose algorithm is used to detect the joint feature points in the image to obtain the joint detection data;

[0055] According to the joint detection data, the spatial coordinates of each feature point are obtained, the known joint connection patterns are matched, the connectivity of all estimated joint combinations is evaluated, and the abnormal connectivity structures are removed to generate the joint position data;

[0056] Based on the joint position data, the spatial distribution state of each joint is analyzed, the angular relationship between adjacent joints is judged, and the established human pose patterns are matched to generate the joint position analysis result.

[0057] Specifically, based on the preliminary pose data, first, a convolutional neural network model required by the OpenPose algorithm is pre-trained using a database of 30,000 human body image samples. These sample images cover a variety of heights, body types, and motion ranges and include common postures such as standing, bending, and raising hands. During pre-training, each image is annotated with joint points, and the pixel coordinates and category information of the main positions such as shoulders, elbows, knees, and ankles are recorded. During the training process, the batch size is set to 32, and the initial learning rate is set to 0.001. The network weights are updated using the backpropagation method in each iteration. If the validation set accuracy is found to be lower than 90% in three consecutive iterations, the learning rate is reduced to half of the original value. This 90% threshold is determined through dozens of rounds of preliminary experiments and used as a reference standard for compromising between accuracy and training duration. After all iterations are completed, an independent test set is selected to evaluate the key point localization deviation and calculate the average precision. After confirming that the detection error remains within a small range, the trained model is loaded and used to perform inference on each frame of the obtained preliminary pose data. The possible joint points of the human body in each frame are sorted according to their confidence levels, and reliable points with a confidence level greater than 0.5 are selected. This value of 0.5 is an empirical threshold obtained through multiple cross-validations based on the noise ratio and occlusion situation in various scenarios. Finally, all key points that meet the confidence requirements are retained, and their pixel coordinates and corresponding confidence levels are recorded to obtain joint detection data.

[0058] According to the joint detection data, first, the pixel coordinates and confidence values of each key point are read in sequence and arranged in the order of the main joints of the human body, including positions such as the neck, shoulders, elbows, wrists, torso, hips, knees, and ankles. Then, the Euclidean distance in the image plane is calculated for every two adjacent joint points and denoted as d ij , if d ij exceeds 100 pixels, it is determined that the distance between the joint pair is too large. This threshold of 100 pixels is an empirical range obtained by referring to the images of 50 testers with a height of about 1.7 meters and considering the image scaling ratio. If there are multiple joint pairs that exceed this threshold simultaneously, these discrete points are ignored and not included in the subsequent connectivity analysis. Subsequently, according to the known human bone connection structure, it is checked whether adjacent joints meet the conventional direction constraints. For example, if the angle formed by the shoulder to the elbow and the elbow to the wrist exceeds 160 degrees, it is regarded as an abnormal connection. This 160 degrees is a threshold obtained by sampling real cases and taking the average extreme actions. When the same joint group is marked as abnormal multiple times, it is excluded as a whole. Then, the remaining normal joint pairs are mapped and the coordinate indices of all joints are recorded to generate joint position data.

[0059] Based on the joint position data, first define the pose vectors for each joint in three-dimensional space and sequentially extract the included angles of vectors from the shoulder to the hip, from the hip to the knee, from the knee to the ankle, etc. If it is detected that the included angle of a certain joint is greater than 150 degrees, it indicates that the joint is close to being straight. This 150 degrees is a reference standard obtained by statistically averaging the joint points of 50 subjects performing different actions such as standing, sitting, and lying down and then adding a floating amount. Then, statistically analyze the relative positions of each joint in space and compare them with an existing human pose library. The human pose library contains 200 annotated samples covering various poses such as standing, sitting, and squatting. When the included angle differences between consecutive joint vectors and a certain group of samples in the library are all less than 10 degrees, it is considered a high match. This value of 10 degrees is a fixed threshold set by summarizing the measurement results of simulated actions. If multiple poses simultaneously meet the threshold standard, they are sorted according to the size of the total angle difference and the matching scheme with the smallest difference is selected. Subsequently, integrate all the matched pose classification information and the coordinates and angle records of adjacent joints to generate the joint position analysis result.

[0060] The steps for obtaining the behavior pattern sequence are as follows:

[0061] Based on the joint position analysis result, extract the time series data of all joints, arrange the spatial coordinates of each joint in chronological order, smooth the joint movements in consecutive frames, and generate the joint movement trajectory data;

[0062] According to the joint movement trajectory data, calculate the curvature change rate of the joint trajectory. The calculation formula is:

[0063]

[0064] where S is the curvature change rate, Q j is the position coordinate of the joint at time T j and k is the total number of frames in the time series.

[0065] Based on the curvature change rate, analyze the time series movement patterns of all joints, classify them according to the change trend of the trajectory, and form the behavior pattern sequence.

[0066] Specifically, based on the joint position analysis results, first, a time series list is created according to each joint coordinate information recorded in the joint position analysis results, and alignment processing is performed in chronological order. The coordinates of the same joint in consecutive frames are matched one by one according to the frame sequence index and written into the same joint trajectory sequence. Then, a smoothing operation is performed on each joint trajectory sequence. The smoothing method includes selecting five consecutive frames as a sliding window in the time dimension and using weighted average calculation. Different weight distributions are set for the current frame and the two adjacent frames before and after. For example, the weight of the current frame is set to 0.4, the two adjacent frames are each 0.3, and the two further adjacent frames are both 0.0 to ignore the influence of too distant frames. The setting of these weights is based on the observation results of the noise distribution during fifty motion capture processes. If the Euclidean distance between frames exceeds 15 pixels twice in a row, this section is regarded as a large displacement area and marked separately. This value of 15 pixels is a reference threshold confirmed by repeatedly collecting human joint motion tests and finding the balance between smoothing and over-smoothing when the general motion distance is no more than 100 pixels. Immediately afterwards, additional smoothing calculation is performed on the trajectory segments marked as large displacements, and quadratic interpolation is performed on the farthest point within each sliding window. The quadratic interpolation process uses a method based on local polynomial fitting to reduce the fluctuations caused by instantaneous displacement mutations. During this process, a range comparison is performed according to the sampling frequency of the motion capture device and the joint movement range. For example, by comparing the displacement amplitude of a joint with the interval from 0 pixel to 100 pixels, if the coordinate increment of a certain joint exceeds the set upper limit of the interval, the point is automatically corrected to a position no more than 5 pixels away from the previous coordinate during quadratic interpolation. This 5-pixel value is the empirical upper bound determined by the maximum smooth displacement of the joint by the tester in a short time. Finally, smoothing processing is completed for all joints, and the joint trajectories in each frame are encoded and recorded to obtain joint motion trajectory data.

[0067] The advantage of the formula is that by combining the second difference component of each joint coordinate in the time series with the square of the time interval, it can intuitively reveal the comprehensive curvature fluctuation of the joint motion trajectory at fine position changes, and quantify the deformation degree experienced by the overall motion process through cumulative summation.

[0068] Q j The obtaining steps of parameter: are that, Q j represents the position coordinate of the joint at time T j When capturing the motion, it is necessary to deploy several infrared cameras to capture the joint reflection markers or determine the joint pixel coordinates through image recognition methods and convert them into spatial coordinates at the action capture stage. Subsequently, the obtained coordinates are recorded according to the frame sequence and arranged in chronological order to form {Q0, Q1,..., Q k}.

[0069] T j The obtaining steps of parameter: are that, Tj Used to identify the acquisition moment of the joint position on the time axis, which can be directly determined by the sampling frequency or video frame rate. For example, if the camera captures at 50 frames per second, the time interval between adjacent frames is 0.02 seconds. The cumulative method starts from the initial moment and accumulates sequentially to form {T0, T1,..., T k}, if the time of the 0th frame is 0 when recording, then at the moment of the 1st frame T1 = 0.02, at the moment of the 2nd frame T2 = 0.04, etc. Finally, by summarizing the time information of all frames, a complete {T j} sequence can be obtained.

[0070] Steps to obtain the k parameter: k represents the total number of frames in the time series, usually determined by multiplying the actual shooting duration of the device by the frame rate. For example, when capturing pose data at 50 frames per second for a 2-second capture duration, 100 frames can be obtained, so k is 100. If shooting continuously for 5 seconds, then k is 250.

[0071] Calculation process: Let k = 3, that is, there are a total of 4 nodes {Q0, Q1, Q2, Q3} and corresponding moments {T0, T1, T2, T3} in the time series. The collected joint coordinates are set as Q0 = (10, 10), Q1 = (12, 11), Q2 = (15, 15), Q3 = (17, 17), and at the same time T0 = 0.00 seconds, T1 = 0.02 seconds, T2 = 0.04 seconds, T3 = 0.06 seconds. First, calculate |Q2 - 2Q1 + Q0| when j = 1, where Q2 - 2Q1 + Q0 = (15 - 24 + 10, 15 - 22 + 10) = (1, 3), and its Euclidean norm is , the time difference (T2 - T1) 2 = (0.04 - 0.02) 2 = 0.0004, so Then calculate |Q3 - 2Q2 + Q1| when j = 2, where Q3 - 2Q2 + Q1 = (17 - 30 + 12, 17 - 30 + 11) = (-1, -2), and its norm is The time difference (T3 - T2) 2 = (0.06 - 0.04) 2 = 0.0004, so Adding the two together gives S ≈ 7905.75 + 5590.25 = 13496.0.

[0072] This result shows that the value of the curvature change rate S is approximately 13496.0. When S is relatively large, it indicates that the joint movement trajectory has significant bending or drastic changes during this time period. Based on this, the activity of the corresponding joint can be judged and further classification or monitoring can be carried out.

[0073] Based on the curvature change rate data obtained previously, first read the curvature change records in the joint movement trajectory in chronological order and create a time series index table to indicate the movement states of different joints in each time period. Then, perform continuous frame scanning on each joint according to this index table. During the scanning process, compare the change amplitude of the joint coordinates one by one and verify it with a pre-set reference range. For example, the reference range can be set between 0 pixels and 80 pixels. When the displacement value of a certain adjacent frame exceeds 80 pixels, mark this section of movement as a high-intensity action. Here, 80 pixels is the upper limit of the average human activity distance obtained by capturing the daily actions of fifty subjects and statistical analysis. This marking process can accumulate multiple high-intensity action segments. If the continuous frame number of some of these segments reaches or exceeds 15 frames, add a separate identifier in the index table and record the specific time interval. The setting of 15 frames also refers to the actual movement experiment results. When the continuous violent movement duration of the subject is above 0.3 seconds (15 frames × 0.02 seconds / frame), significant acceleration fluctuations often occur. Next, combine the above-marked trajectory information with the curvature change rate, conduct segmented statistics on the curvature value sequence of each joint in the high-intensity action segment, and distinguish between two categories: moderate bending in the range of 30 to 60 and high bending above 60. The two values of 30 and 60 are based on the observation results of multiple batches (at least ten batches) of human joint movement experiments. When the subjects are walking normally or standing, the curvature is usually below 30, while when quickly turning or bending the joints, the curvature is greater than 60. If the curvature value of a certain segment is higher than 60 in most frames, classify this segment into the "extreme bending movement" list, and focus on tracking the coincidence degree between this list and the high-intensity action segment in the subsequent analysis. If the coincidence degree is too high, it means that both the movement amplitude and bending degree of the joint are in the upper limit area during this time range. After completing all the statistics and classification, finally mark and sort out various types of movement segments and arrange them in chronological order, and then compare them with the previously obtained curvature change rate to obtain the comprehensive movement mode characteristics of each joint at different time periods, and form a behavior mode sequence accordingly.

[0074] The steps to obtain the posture classification result are as follows:

[0075] Based on the behavior mode sequence, extract the joint movement states within all time segments, traverse all time windows to calculate the relative displacement vectors of the joints, and calculate the relative rotation matrices of the joint points based on the three-dimensional Euclidean space transformation to determine the angle change trend between each time point and generate behavior feature data;

[0076] According to the behavior feature data, calculate the posture classification score. The calculation formula is:

[0077]

[0078] Among them, F is the posture classification score, As is the joint angle change rate at time s, B s is the joint acceleration change value at time s, C s is the Euclidean distance between adjacent joints at time s, D s is the trajectory curvature at time s, E s is the inertia change rate at time s, and p is the total number of frames within the time window;

[0079] Based on the posture classification score, compare with the known posture model database, and combine with the set of motion patterns to analyze the classification characteristics of different posture states, match the states of standing, sitting / lying, or falling, and generate the posture classification result.

[0080] Specifically, based on the behavior pattern sequence, select the time segments of interest to extract the joint motion state data, and perform grouped statistics in the order of frames within each time segment. Each group contains ten frames of data and corresponds to a sampling period of 0.2 seconds. Since the capture device is set to a frame rate of 50 frames per second, at this time, sequentially check key parts such as the shoulder joint, elbow joint, and knee joint in each group. Calculate the three-dimensional coordinate difference between adjacent frames and convert it into a vector representation, then project the vector onto the corresponding initial vector of the bone along the coordinate axes. Furthermore, extract the vector angle in the three-dimensional space as the rotation reference, and then construct a relative rotation matrix accordingly. If it is found that the rotation angle of a certain joint is higher than 180 degrees in both adjacent groups, then this joint is regarded as having undergone a large twist. This 180 degrees comes from the statistical value of the normal joint activity limit in clinical anatomy. At the same time, an additional record is made for this time segment to indicate that the joint activity appears in the limit state. After the recording is completed, read the relative rotation matrices of all time windows and compare the changes in the elements on the main diagonal of the matrix one by one. These elements can be used to describe the adaptive adjustment of the rotation matrix in the rigid body change. If the elements on the main diagonal of the matrix exceed a preset threshold (for example, exceed 1.5 or less than -1.5, and the specific value comes from the test results of bone simulation in multiple scenarios) continuously twice, then this time window is listed as an abnormal activity stage and an additional mark is made. After completing the overall marking work, summarize the angle evolution information of all joints, and compare and map it with the behavior pattern sequence. For example, if a certain time period has been marked as a high-intensity action and the elements on the main diagonal of its rotation matrix also show multiple jumps, it can be judged that the joint motion is changing rapidly. Finally, summarize the angle change trend between each time point to generate the behavior feature data.

[0081] The advantage of the formula is that the difference between the joint angle change rate and the joint acceleration change value within all time frames (multiplied by the Euclidean distance between adjacent joints) is accumulated as the numerator, and then the sum of the trajectory curvature and the inertia change rate plus 1 is used as the denominator for the frame-level summation, so as to comprehensively consider the impacts of angle difference, spatial distance, curvature, and inertia accumulation on the final pose score, and balance data with different dimensions in the form of cube root, making the score more discriminative and stable.

[0082] A s The acquisition steps of the parameter are as follows: A s represents the joint angle change rate at time s. First, the angle difference needs to be calculated based on the bone connection angles at each time point in the previously obtained behavior pattern sequence, and then divided by the time interval between adjacent frames to obtain the angle change rate. Each s is derived from actual pose capture. The angle difference can usually be obtained by comparing the bone angles at the previous moment and the current moment. If the capture device frame rate is 50 frames per second, the time interval for each frame is 0.02 seconds. For example, when the shoulder joint angle increases from 60 degrees in the previous frame to 64 degrees in this frame, the angle change amount is 4 degrees, and the angle change rate is 4 / 0.02 = 200 degrees per second. Processing all frames in chronological order can form {A1, A2,..., A p}, for example, the angle change rates collected in 10 consecutive frames are several values such as 200, 220, 250, etc.

[0083] B s The acquisition steps of the parameter are as follows: B s is the joint acceleration change value at time s, which needs to be obtained based on the difference of s A. The ΔA s can be calculated between every two adjacent moments, where s ΔA s-1 = A s - A s and then divided by the 0.02 seconds between frames to obtain B s . These data can be derived from continuous pose records. If it rises from 200 degrees per second to 220 degrees per second at a certain moment, then ΔA 2 = 20, and ΔA p / 0.02 = 1000 degrees per second

[0084] C s The acquisition steps of the parameter are as follows: C s is used to represent the Euclidean distance between adjacent joints, and the distance formula can be applied in three-dimensional coordinates where (x1, y1, z1) and (x2, y2, z2) are the spatial positioning data of adjacent bone nodes. The distances calculated for each frame are uniformly numbered, resulting in {C1, C2,..., C p}, for example, if the detected distance between the shoulder and the elbow is approximately 0.25 meters at a certain moment and becomes 0.30 meters at another moment, then 0.25 or 0.30 can be recorded in the corresponding C s .

[0085] D s The steps for obtaining the D parameter are as follows. D s represents the trajectory curvature value at time s. The curvature magnitude can be indexed and extracted frame by frame based on the previously obtained trajectory bending analysis results. For example, when calculating the curvature change rate, a local curvature quantity is produced for each frame. Aligning this value with the frame sequence forms {D1, D2,..., D p}, and the values generally range from about 10 to 150, depending on the amplitude and degree of bending of the action.

[0086] E s The steps for obtaining the E parameter are as follows. E s represents the inertial change rate, which needs to be obtained by integrating the acceleration within a certain interval based on the acceleration information. When the sensor captures data, the acceleration within 0.02 seconds can be set for each acquisition and integrated. If the integrated value fluctuates significantly, it is considered that the inertial change rate is high, and it is recorded at the corresponding moment and summarized into {E1, E2,..., E p}

[0087] The steps for obtaining the p parameter are as follows. p represents the total number of frames in the time window. Usually, when analyzing behavioral characteristics, several frames are combined into a statistical period, such as 10 frames or 20 frames, etc. If 10 frames are combined, then p = 10.

[0088] Calculation process:

[0089] Let p = 5, and successively give A1 = 210, A2 = 250, A3 = 260, A4 = 280, A5 = 300 degrees / second, B1 = 0, B2 = 2000, B3 = 500, B4 = 1000, B5 = 0 degrees / second, C1 = 0.24, C2 = 0.25, C3 = 0.26, C4 = 0.27, C5 = 0.28 meters, D1 = 20, D2 = 40, D3 = 60, D4 = 80, D5 = 100, E1 = 0.12, E2 = 0.13, E3 = 0.15, E4 = 0.16, E5 = 0.17. Then the numerator is:

[0090]

[0091] The denominator is:

[0092]

[0093] Divide the numerator by the denominator to approximately 828.7 / 305.73 ≈ 2.71, and then take the cube root. Therefore, F = 1.40.

[0094] This result indicates that the attitude classification score for the current time window is approximately 1.40. When combined with the subsequent determination rules, when F is between 1.0 and 2.0, it can be considered that the movement amplitude is medium. If it exceeds 2.5, it indicates that the movement amplitude is significant. Furthermore, based on this, states such as standing, sitting / lying, or falling can be matched in the attitude model database, thereby completing the corresponding attitude classification steps.

[0095] Based on the attitude classification score, compare it with a set of local registered human attitude model databases. For each attitude in the database, a corresponding score range is configured and distinguished by numbers. For example, a score range of 1.0 to 2.0 is configured for the standing attitude, a score range of 2.0 to 3.0 is configured for the sitting / lying attitude, and a score range exceeding 3.0 is configured for the falling attitude. Subsequently, compare the attitude score of each time window calculated previously with these ranges and record the attitude label corresponding to this window. Then, further integrate the action categories of each time window with their attitude labels according to the aggregated motion pattern set, and merge and mark the segments with continuous identical attitude labels on the time axis. If a certain attitude label continuously appears in several windows and is not covered by a higher score range, then this attitude segment is stably marked. If it is detected that the score of a certain period spans multiple range intervals, then conduct repeated verification and search for the most likely attitude model in the database. After confirming that there are no other abnormalities, record this segment as the corresponding attitude category. Finally, extract the occurrence times, continuous frame numbers, and start and end times of each period attitude during the entire action capture process, and arrange these marks in order and indicate the corresponding attitude information to generate the attitude classification result.

[0096] The preset acquisition steps for the trigger decision are as follows:

[0097] Based on the attitude classification result, extract the attitude state data for all time periods, analyze the distribution of different attitudes in the time series, calculate the duration of each attitude, and conduct segmented marking on the continuously changing states. Combine the indoor environment parameters to extract attitude anomaly data and generate environment adjustment requirement data;

[0098] According to the environment adjustment requirement data, calculate the environment adjustment decision score. The calculation formula is:

[0099]

[0100] Among them, V is the environment adjustment decision score, P c is the duration of the current attitude, P dis the duration of the previous posture, CQ x is the current light intensity, CQ y is the last recorded light intensity, R is the rate of change of indoor temperature, S is the rate of change of humidity, and CT is the preset response time interval;

[0101] Based on the environmental adjustment decision score, decide whether to adjust the lighting brightness, air conditioning temperature or curtain status, and generate a preset trigger decision.

[0102] Specifically, based on the posture classification results, the start and end time of each posture in all the recorded posture state data is first read and converted into a time period form, and then the cumulative duration of the same posture in adjacent time periods is counted in chronological order, and the segments with a cumulative duration greater than 5 seconds are merged and marked. The value of 5 seconds comes from the actual observation of the time required for posture transitions of 20 subjects in daily activity scenarios and the average threshold value obtained by statistics. Then, these merged posture time periods are subjected to transition detection between adjacent postures and whether frequent posture switching occurs within a limited time. For example, when the number of consecutive switching reaches 3 times and the total time interval does not exceed 10 seconds, it is regarded as a continuously changing state. The 10 seconds is taken from the short-term rapid conversion threshold commonly used in motion monitoring. After multiple data comparisons, it is found that it can better identify frequently changing posture segments. Subsequently, these continuously changing states are segmented and marked in the record. Up-conversion identification, when identifying, the degree of high-frequency switching is prompted by comparing the difference in duration between the previous posture and the current posture, and comparing the duration difference with the preset minimum stable duration of 2 seconds. The 2-second threshold is also obtained based on real scene statistics. When the duration difference is less than 2 seconds and the switching frequency is greater than 3 times, it is included in the abnormal state list, and then the posture abnormality is identified in combination with the indoor environmental parameters, such as temperature fluctuations between 24°C and 28°C, humidity fluctuations between 40% and 60%, and light intensity changes between 300lx and 600lx. If the posture abnormality occurs when the current temperature or humidity is close to the upper limit of the set safety range, the abnormal information will be separately marked and stored together with the corresponding timestamps of the posture state change period and the environmental indicators. Finally, all abnormal marks are classified and merged, and the statistical abnormal segment duration and environmental condition indicator values are referenced to form environmental adjustment demand data.

[0103] The benefit of the formula is that it incorporates the difference in duration between the current posture and the previous posture, as well as the light intensity, temperature change rate and humidity change rate into a comprehensive metric, and normalizes the comprehensive metric through the preset response time interval in the denominator, thereby taking into account the dynamic transformation of human posture and the real-time fluctuation of environmental conditions, and can more sensitively determine whether environmental adjustment operations are needed.

[0104] P cParameter: The acquisition step of P is as follows c represents the duration of the current pose, which needs to be obtained by extracting the cumulative duration of the same pose within a continuous time period from the previously summarized pose state data. First, obtain the start and end frame indices of the identified pose label on the time axis, convert the number of frames to seconds through the frame rate, and then add up the periods of multiple identical poses without being inserted by other poses. For example, if it is recorded that the standing pose appears from frame 100 to frame 150 and the frame rate is 25 frames per second, then the duration of this segment is (150 - 100) / 25 = 2 seconds. If the same pose appears again from frame 151 to frame 175 later, then calculate its duration again, approximately (175 - 151) / 25 = 0.96 seconds. After summing up, the duration P of the current pose is obtained c = 2.96 seconds.

[0105] P d Parameter: The acquisition step of P is as follows d represents the duration of the previous pose. It is necessary to record the start and end times of the previous pose before the current pose is switched and calculate its total duration in the same way according to the number of frames and the frame rate. When it is found that the pose switches from A to B, the duration of A is archived and defined as P d , for example, if the previous pose A lasts from frame 80 to frame 99, and the same sampling rate of 25 frames per second is used, then P d = (99 - 80) / 25 = 0.76 seconds.

[0106] CQ x Parameter: The acquisition step of CQ is as follows x represents the current light intensity, which needs to be read through the brightness data obtained in real-time by the indoor photometric sensor during pose detection and take its average value to represent the current ambient light level. For example, by collecting the light intensity data stream within 500 milliseconds and averaging all the samples in it. If a total of 50 collections are made and each reading is between 280 lx and 310 lx, then add up these readings and divide by 50 to get approximately 295 lx. At this time, CQ x = 295 lx.

[0107] CQ y Parameter: The acquisition step of CQ is as follows y represents the previously recorded light intensity, usually from the photometric sensor measurement value saved in real-time at the end of the previous time period pose. Similar to CQ x it is also obtained by averaging the light measurement data over a period of time. When the pose is switched, the light intensity average value at this moment is set as CQ y for reading during the next operation. For example, when the system measures and records the light value of about 310 lx at the end of the previous pose, then CQ y = 310 lx.

[0108] For the R parameter: The acquisition steps are as follows. R represents the change rate of the indoor temperature. It is necessary to measure the current indoor temperature value multiple times within a time interval and convert its increase or decrease rate into the change amplitude per second or per minute. For example, measure the temperature once every 5 seconds with a temperature sensor and record the results of 10 consecutive samplings. Then calculate the difference between the previous and current measurement values and divide it by the measurement time. The final value of R can be obtained by taking the average of the change rates calculated multiple times. If it is observed that the temperature rises between 24.0°C and 24.5°C during 10 measurements and the total time taken is about 50 seconds, then the total increase is 0.5°C, and the average increase per second is 0.5 / 50 = 0.01°C. Thus, R = 0.01°C / second.

[0109] For the S parameter: The acquisition steps are as follows. S represents the change rate of humidity. Similar to temperature, the humidity sensor can record multiple sampling readings within the same time step and divide the difference between the previous and current sampling values by the length of the sampling interval to calculate the increase or decrease in humidity per second or per minute. If 10 consecutive data points are collected and it is recorded that the humidity changes from 52%RH to 55%RH in 40 seconds, then the total change is 3%RH, and the average change per second is 3 / 40 = 0.075%. Therefore, S = 0.075% / second.

[0110] For the CT parameter: The acquisition steps are as follows. CT represents the preset response time interval. In actual operation, a fixed value can be set in the control system so that the system will perform a complete environmental decision-making operation every CT seconds. This value is usually set according to user requirements and the reaction ability of the device. For example, in hot regions where the air conditioner has a high cooling frequency, CT can be set to 10 seconds to increase the update frequency of the decision-making. If in a region with relatively stable climate, CT can be appropriately extended to 30 seconds to allow the system to adjust the environment more smoothly. The specific value needs to be selected by considering multiple factors such as environmental monitoring and human comfort assessment. For example, if it is found through testing that CT = 15 seconds can balance energy conservation and real-time performance, then this value is adopted.

[0111] Calculation process: Let P c = 3.2 seconds, P d = 2.4 seconds, CQ x = 350 lx, CQ y = 310 lx, R = 0.02°C / second, S = 0.03% / second, CT = 10. Among these, these parameters are all from the results of the previous posture duration statistics and indoor sensor monitoring. Substitute them into the formula in turn:

[0112] |P c -P d | = |3.2 - 2.4| = 0.8

[0113]

[0114] Add the two parts together:

[0115] 0.8 + 40.0000 = 40.8

[0116] At the denominator:

[0117] CT + 1 = 10 + 1 = 11

[0118] Finally, we get:

[0119]

[0120] This result indicates that the current environmental adjustment decision score is approximately 3.71. When this value is close to or exceeds 3, it means that the duration of the current posture is significantly different from the previous one, and the combined changes in light, temperature, and humidity are significant. More consideration needs to be given to the triggering intensity or range of adjustment actions.

[0121] Based on the environmental adjustment decision score, first compare the calculated score value with the pre - determined threshold range. For example, when the score exceeds 3, it is marked as a high - demand adjustment area; if it is between 2 and 3, it is marked as a medium - demand area; if it is below 2, it is marked as a low - demand area. These threshold values are three demarcation points summarized based on hundreds of user preference surveys and environmental debugging results, and are encoded in the system for automatically judging the triggering level of specific adjustment actions. Subsequently, when the current score value is greater than 3, read the switches and power settings of the involved lighting, air - conditioning, and curtain devices and compare them with parameters such as indoor temperature, humidity, and light intensity. If the detected light is lower than 300 lx, send a command to increase the brightness of the lighting device; if the monitored indoor temperature rise rate continues to be higher than 0.02 °C / second, execute the air - conditioning cooling setting and lower the target temperature by 1 °C to 2 °C. At the same time, check the curtain opening degree in the previous time period and combine it with the current sunlight intensity. When the sunlight is strong and the indoor light has reached more than 400 lx, automatically close the curtain to 50% opening to avoid excessive light entering. Finally, summarize the setting values and execution times of each device and uniformly send control commands to generate a preset trigger decision.

[0122] The steps to obtain the smart home response status are as follows:

[0123] Based on the preset trigger decision, parse the home environment adjustment command, extract the target device, adjustment parameters, and execution timing in the control command, convert the command into an instruction format recognizable by the device, and generate controller instruction data;

[0124] According to the controller instruction data, send instructions to the home smart devices, execute the adjustment of lighting brightness, air - conditioning temperature, and curtain status, monitor the execution feedback of the devices, verify the correct transmission and execution of the instructions, and generate device response data;

[0125] Based on the device response data, analyze the operating status of home smart devices, determine whether the adjustment meets the preset requirements, record the status data of the device after adjustment, and resend instructions to devices that have not been executed or responded abnormally to form a smart home response status.

[0126] Specifically, based on the preset trigger decision, the target device identification and the corresponding adjustment instruction set are first read from the statistical environmental adjustment demand data, and the instruction type and parameter value of each specific command in the instruction set are parsed. If the instruction type belongs to light brightness adjustment, the specified brightness level is read and it is determined whether it is in the range of 0 to 100%. If the brightness level exceeds 100%, it is merged and corrected to a maximum limit not higher than 100%. This 100% value is determined based on the power limit of the lamp manufacturer under long-term working conditions and is marked as a 100% brightness upper limit at the factory. If the instruction type is temperature adjustment, the target temperature parameter is further read and compared with the established indoor acceptable temperature range of 18°C to 30°C. If the target temperature is lower than 18°C or higher than 30°C, it is corrected to the upper and lower limits of the range. This range is a commonly used adjustment range based on the comfort temperature experimental results of most users in residential areas. If the instruction type is curtain state control, the curtain opening degree is obtained from the instruction set and compared with the pre-established 0% to 100% opening range. If it exceeds this range, Set the opening degree to the nearest end value. After completing the confirmation of all command parameters, sort each command according to its execution sequence. If it is detected that there are multiple commands to be executed within the same second, they are arranged according to the device priority. The device priority is obtained by querying the pre-established priority table. For example, air conditioning adjustment is listed as a higher priority than curtain opening operation, and curtain opening operation takes precedence over light brightness fine-tuning. This priority table comes from the investigation of actual user needs and the research on equipment circuit load distribution. When the instructions are sorted, the instruction type, the corrected parameter value and the execution sequence are integrated and assembled into an instruction data format that can be parsed by the corresponding device. For example, lighting equipment requires receiving instructions in a fixed structure such as "function code + brightness value + check code", while air conditioning equipment requires "function code + mode + target temperature + check value" format for identification. Curtain equipment also has corresponding format definitions, and needs to be accompanied by information such as the safety delay of the curtain motor. All assembled instruction data are uniformly labeled and distinguished from different devices with corresponding identifiers. After the assembly is completed, the controller instruction data is formed.

[0127] According to the controller instruction data, each device command and its execution timing are read one by one, and this timing is compared with the current system time. If the time difference between the system time and the set start time of the command is within 0.1 second, it is considered that the command has reached the execution time. Otherwise, for the commands that have not reached the time, wait until the timing is satisfied before execution. Then, the instructions that match the execution time are sent to the corresponding home intelligent devices according to the device type. For example, the brightness adjustment instruction is sent to the lighting device via wireless signal and the returned value of the lighting brightness is received in real time. The temperature adjustment instruction is sent to the air conditioner and the current operating information such as the outlet air temperature and wind speed status of the air conditioner is periodically queried. The curtain opening degree adjustment instruction is sent to the curtain driving motor and at the same time, it is monitored whether there is an overload of the motor current. In the judgment process, the current can be compared with the preset range between 0A and 5A. If the current exceeds 5A, the corresponding alarm information is immediately recorded and the next operation of this device is paused. At any time, it is checked whether such an abnormality is self-recovered or requires manual intervention. For the instructions that are normally executed, the device status is queried at a frequency of every 1 second, and information such as the actual brightness value of the light, the current temperature difference of the air conditioner, and the curtain opening percentage is recorded. The 1-second query frequency here is a compromise frequency selected after repeated measurement in the test to reduce the network and device load. When the difference between the value returned by the device and the instruction target value exceeds 2%, the same instruction is sent again for correction. This 2% difference is a relatively balanced accuracy standard during actual maintenance. After waiting for the device to return values multiple times and showing a trend towards the target value, the results are summarized based on the final measured data and the system timestamp to generate device response data.

[0128] Based on the device response data, first, check one by one from the recorded final return values and execution statuses of each device whether the actual execution results meet the requirements of the corresponding instructions. For example, if the difference between the returned value of the light brightness and the brightness required by the command is greater than 5%, it is considered to exceed the normal deviation range. This 5% threshold is generated based on the average value of the brightness attenuation and control accuracy tests of ten lamps of the same model after long-term operation, and it has been verified in multiple experiments that it can effectively distinguish normal illuminance deviation from abnormal control. For air conditioning equipment, compare whether the difference between the target temperature and the returned air outlet temperature is within 0.5°C. This 0.5°C range is obtained through the refrigeration system data provided by the manufacturer and the calculation during the operation of the equipment power. For curtain equipment, sequentially read the returned opening percentage and compare it with the command value. If the difference exceeds 3% or continuous overload alarm occurs, it is determined as an abnormal state, and a process of resending the command for this equipment is arranged in subsequent commands. Subsequently, summarize the device operation information that has met the instruction requirements in chronological order. If a situation where a large area of parameter deviation or invalid instruction occurs for a certain device, screen out the corresponding time period of this device and resend the instruction. When sending, still maintain the original parameter values and execution timing and add a reissue mark. If it still cannot respond again, record it in the list that requires manual diagnosis and perform additional processing. Finally, record the operation status and final output values of all devices that have ended normally or completed calibration, and uniformly count to form a complete smart home response status.

[0129] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A smart home scenario interaction method based on image recognition technology, characterized in that Including the following steps: Based on the real-time image data of the occupant captured by the camera, extract the pixel points and color depth information in the image to obtain preliminary pose data; Based on the preliminary pose data, apply the OpenPose algorithm to detect the joint positions, identify the pose of the occupant, and generate the joint position analysis result; Based on the joint position analysis result, perform serialization processing on the motion trajectory of each joint through time series analysis to obtain the behavior pattern sequence; based on the behavior pattern sequence, analyze the behavior state of the occupant, where the behavior state includes standing, sitting / lying, or falling, and generate the pose classification result; Based on the pose classification result, decide whether it is necessary to adjust the home environment, including lighting or air conditioning adjustment, and generate a preset trigger decision; Based on the preset trigger decision, send an adjustment command to the home intelligent device through the controller, including the settings of indoor lighting brightness, air conditioning temperature, and curtain status, to obtain the smart home response status.

2. The smart home scenario interaction method based on image recognition technology according to claim 1, wherein, The steps for obtaining the preliminary pose data are as follows: Capture the real-time image data of the occupant through the camera, perform color and brightness correction on each frame of the image to obtain the corrected real-time video data; Based on the corrected real-time video data, apply bicubic interpolation to enhance the resolution of each frame of the image to obtain the preliminary pose data.

3. The smart home scenario interaction method based on image recognition technology according to claim 1, wherein The steps for obtaining the joint position analysis result are as follows: Based on the preliminary pose data, use the OpenPose algorithm to detect the joint feature points in the image to obtain the joint detection data; According to the joint detection data, obtain the spatial coordinates of each feature point, match the known joint connection patterns, perform connectivity evaluation on all estimated joint combinations, and eliminate abnormal connectivity structures to generate the joint position data; Based on the joint position data, analyze the spatial distribution state of each joint, judge the angular relationship between adjacent joints, match the established human pose patterns, and generate the joint position analysis result.

4. The smart home scenario interaction method based on image recognition technology according to claim 1, wherein, The steps for obtaining the behavior pattern sequence are as follows: Based on the joint position analysis result, extract the time series data of all joints, arrange the spatial coordinates of each joint in chronological order, smooth the joint motion in consecutive frames, and generate the joint motion trajectory data; According to the joint motion trajectory data, calculate the curvature change rate of the joint trajectory, and the calculation formula is: where S is the curvature change rate, Q j is the position coordinate of the joint at time T j moment, and k is the total number of frames in the time series; Based on the curvature change rate, analyze the time series motion patterns of all joints, classify according to the change trend of the trajectory, and form the behavior pattern sequence.

5. The smart home scenario interaction method based on image recognition technology according to claim 1, wherein The steps for obtaining the pose classification result are as follows: Based on the behavior pattern sequence, extract the joint motion states within all time segments, traverse all time windows to calculate the relative displacement vector of the joints, and calculate the relative rotation matrix of the joint points based on the three-dimensional Euclidean space transformation to determine the angle change trend between each time point, and generate the behavior feature data; According to the behavior feature data, calculate the pose classification score, and the calculation formula is: Among them, F is the posture classification score, A s is the joint angle change rate at time s, B s is the joint acceleration change value at time s, C s is the Euclidean distance between adjacent joints at time s, D s is the trajectory curvature at time s, E s is the inertia change rate at time s, and p is the total number of frames within the time window; Based on the pose classification score, compare with the known pose model database, combine the motion pattern set to analyze the classification features of different pose states, match the states of standing, sitting / lying, or falling, and generate the pose classification result.

6. The smart home scenario interaction method based on image recognition technology according to claim 1, wherein The steps for obtaining the preset trigger decision are as follows: Based on the posture classification result, extract the posture state data in all time periods, analyze the distribution of different postures in the time series, calculate the duration of each posture, segment and mark the continuously changing states, extract posture anomaly data in combination with indoor environmental parameters, and generate environmental adjustment requirement data; According to the environmental adjustment requirement data, calculate the environmental adjustment decision score. The calculation formula is: Among them, V is the environmental adjustment decision score, P c is the duration of the current posture, P d is the duration of the previous posture, CQ x is the current light intensity, CQ y is the previously recorded light intensity, R is the change rate of the indoor temperature, S is the change rate of the humidity, and CT is the preset response time interval; Based on the environmental adjustment decision score, decide whether to execute the adjustment of the light brightness, air conditioner temperature or curtain state, and generate a preset trigger decision.

7. The smart home scenario interaction method based on image recognition technology according to claim 1, wherein, The steps for obtaining the smart home response status are as follows: Based on the preset trigger decision, parse the home environment adjustment command, extract the target device, adjustment parameters and execution timing in the control instruction, convert the command into an instruction format recognizable by the device, and generate controller instruction data; According to the controller instruction data, send instructions to the home smart devices to execute the adjustment of the light brightness, air conditioner temperature and curtain state, monitor the execution feedback of the devices, verify the correct transmission and execution of the instructions, and generate device response data.

8. The smart home scenario interaction method based on image recognition technology according to claim 7, characterized in that, The steps for obtaining the smart home response status further include: Based on the device response data, analyze the operating status of the home smart devices, judge whether the adjustment meets the preset requirements, record the state data of the devices after adjustment, and resend instructions to the devices that have not been executed or have abnormal responses to form the smart home response status.

Citation Information

Cited By

  • Campus sports digital management method and system

    CN120564274A

  • Campus sports digital management method and system

    CN120564274B

  • Motion posture recognition method and system based on deep learning

    CN121305683A