Identification methods and electronic devices

By extracting behavioral semantic feature data and combining it with a multi-sensor sliding window segmentation method, the accuracy and efficiency issues of intelligent devices in recognizing human behavior are solved, achieving more efficient and accurate behavior recognition.

CN118057369BActive Publication Date: 2025-10-31HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211459119.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-10-31
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

Existing smart devices have low accuracy and poor recognition efficiency when recognizing human behavior, and cannot effectively utilize sensor data for accurate behavior recognition.

Method used

By extracting behavioral semantic feature data and combining it with data from accelerometer, gyroscope, geomagnetic sensor and WIFI sensor, the system uses sliding window segmentation and multi-sample window segmentation methods to identify user behavior patterns, including behaviors such as being stationary, walking, going up and down stairs, going up and down elevators and escalators.

Benefits of technology

It improves the accuracy and efficiency of human behavior recognition, reduces device power consumption, and maintains the accuracy of recognition results when multiple sensors fail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118057369B_ABST
    Figure CN118057369B_ABST
Patent Text Reader

Abstract

This application provides an identification method and an electronic device. The method extracts behavioral semantic features from data collected by the electronic device's sensors that better reflect user behavior patterns, and combines these behavioral semantic features to identify user behavior, effectively improving the accuracy of user behavior identification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to identification methods and electronic devices. Background Technology

[0002] According to relevant research, people spend more than 80% of their time indoors each day. Therefore, research on indoor behavior is crucial. Accurately identifying people's indoor behavior can significantly advance many research fields. For example, it can study whether people change floors, explore their indoor activities, and even monitor the health of elderly people and children at home, detecting potential falls or other dangerous actions. Currently, with the increasing number of sensors integrated into smartphones and the continuous updates to mobile phones, how to utilize these sensors to obtain people's current movement status is becoming a hot research topic.

[0003] Currently, smart devices require a large amount of data collected by sensors to train relevant predictive models, which are then used to predict the current state of human behavior. However, the accuracy and efficiency of current smart device recognition of human behavior are relatively low. Therefore, there is a need to research more accurate and efficient methods for recognizing human behavior. Summary of the Invention

[0004] The purpose of this application is to provide an identification method and an electronic device. This electronic device can extract behavioral semantic features that better reflect user behavior patterns based on data collected by sensors, and then combine these behavioral semantic features to identify user behavior. Implementing this identification method can effectively improve the accuracy of the electronic device's identification results of user behavior.

[0005] The aforementioned and other objectives will be achieved through the features described in the independent claims. Further implementations are illustrated in the dependent claims, the specification, and the drawings.

[0006] In a first aspect, this application provides an identification method, the method comprising: extracting features from target data to obtain feature data, the feature data including behavioral semantic feature data, the target data being obtained based on initial data collected by sensors in an electronic device, the behavioral semantic feature data including first feature data characterizing the electronic device being in a state of hypergravity / weightlessness, and / or second feature data characterizing the orientation change of the electronic device; performing feature analysis on the feature data to obtain a behavior identification result, the behavior identification result including a first identification result, the first identification characterizing the user's current behavior pattern.

[0007] The target data can be obtained based on initial data collected by sensors in the electronic device. The specific process involved can be found in the descriptions of the foregoing and subsequent embodiments, and will not be repeated here. In this method, the electronic device may include multiple sensors, such as an accelerometer, a gyroscope, a geomagnetic sensor, and a Wi-Fi sensor. These sensors can collect corresponding sensor data at a certain frequency, and this data constitutes the initial data.

[0008] Feature extraction refers to the process of transforming target data into feature data that can characterize user behavior. In this application, in order to focus on studying human behavior patterns and accurately identify five behaviors of the human body: stillness, walking, climbing stairs, escalator climbing, and elevator climbing, the electronic device combines the differences in details and data of various user behavior patterns during the feature extraction process of the target data to extract the behavioral semantic feature data from the target data. This behavioral semantic feature data includes first feature data representing the electronic device being in a state of hypergravity / weightlessness, and / or second feature data representing changes in the orientation of the electronic device (user).

[0009] Specifically, the feature data may also include statistically significant feature data and physically significant temporal feature data. Compared to statistical and temporal feature data, the behavioral semantic features better reflect user behavior patterns. For example, while walking, there is generally no more than one consecutive data frame of weightlessness or weightlessness, but the human body experiences momentary and significant weightlessness or weightlessness. However, when in an elevator, although the human body is mostly stationary relative to the elevator, the elevator's acceleration and deceleration during its movement cause the human body to experience a motion state of slight weightlessness (weightlessness) – constant speed – slight weightlessness (weightlessness), with the weightlessness or weightlessness state typically lasting about 2-4 seconds. Another example is when a person moves from one floor of a staircase to an adjacent floor, requiring a 180° turn. Therefore, this application, by extracting behavioral semantic features that reflect the differences in details and data of different behaviors, can effectively improve the accuracy of user behavior recognition results.

[0010] In conjunction with the first aspect, in one possible implementation, the feature extraction of the target data includes: segmenting the initial data using a sliding window of a first length to obtain first data, the first data including multiple data frames of the first length; segmenting the first data using at least two sample windows to obtain at least two sets of data fragments, wherein the length of any sample window in the at least two sample windows is different from the length of the other windows in the at least two sample windows; and using the at least two sets of data fragments as the target data.

[0011] It should be understood that since actions take a certain amount of time to occur, the data collected by the sensor at a single moment is insufficient to reflect the characteristics of the user's actions. Therefore, after obtaining the initial data, the electronic device can select a sliding window of a certain length and coverage to segment the initial data collected by the sensor, that is, to divide the long-term sequence of initial data into smaller data frames (the length of each data frame is the length of the aforementioned sliding window), thus obtaining the first data, which facilitates feature extraction.

[0012] Furthermore, samples extracted from a single data frame cannot adequately reflect the semantic features before and after the current moment. For example, during elevator operation, there are periods of constant speed. If only a single data frame is used to extract features at this moment, the features will not differ significantly from those at a stationary position. In other words, if only a single data frame is used for feature extraction, the information captured within the window is too limited, potentially leading to lower accuracy in the prediction model's output. Therefore, electronic devices can segment the first data using sample windows based on data frames, extracting features in batches of several data frames during feature extraction. The duration of meaningful feature data varies depending on the behavioral pattern. For instance, when a user is moving between floors in an elevator, a longer sample window is needed to fully extract features indicating weightlessness and g-force. However, when the user is in a stationary or moving mode, only a few data frames are needed to reflect the user's motion characteristics, allowing for a shorter sample window. Therefore, in this embodiment, the electronic device can use the at least two sample windows to segment the first data to obtain the at least two sets of data fragments. Using the at least two sets of data fragments as the target data can make the subsequently obtained feature data more referential and further improve the accuracy of the recognition results.

[0013] In conjunction with the first aspect, in one possible implementation, the step of using at least two sample windows to segment the first data to obtain at least two sets of data fragments includes: using three sample windows to segment the first data to obtain three sets of data fragments, wherein the window lengths corresponding to the three sample windows are 5 times the first length, 10 times the first length, and 15 times the first length, respectively.

[0014] In this embodiment, the at least two sample windows can be three sample windows, and the aforementioned at least two sets of data segments constitute three sets of data segments. Furthermore, the specific sizes of the three sample windows can be set to 5 data frame lengths, 10 data frame lengths, and 15 data frame lengths, respectively. These window sizes were determined through observation of a large amount of experimental data. When the sampling frequency of each sensor in the electronic device is 100Hz, the data frame length is 2.56s (one data frame length), and the number of sample windows is 3, the prediction model achieves the best accuracy when the sample window sizes are set to 5, 10, and 15 data frame lengths, with an output accuracy of approximately 0.98.

[0015] In conjunction with the first aspect, in one possible implementation, the behavior recognition result further includes a second recognition result, which characterizes the point in time when the user switched behavior modes.

[0016] In this embodiment, the behavior recognition result may further include the user's landmark recognition result, i.e., the second recognition result. The landmark recognition result represents the moment when the user switched behavior modes. That is, the landmark point recognition result can be used to reflect the moment when the user's behavior switched in the behavior mode recognition result. For example, if a certain sequence of numbers reflects that the user successively performed a horizontal movement mode and a cross-level movement mode, then the landmark point recognition result can reflect at what moment the user ended the horizontal movement and began the cross-level movement.

[0017] It should be understood that the landmark recognition result is derived by the electronic device based on data collected by sensors to predict the user's behavior. However, there may still be some time error between the predicted and actual landmarks. Therefore, during the training of the prediction model, the electronic device can compare the landmark recognition result with the user's actual landmark data to determine the delay time (number of samples) between the start and end times of the predicted and actual landmarks, the prediction landmark recognition accuracy (the percentage of landmarks with recognition errors ≤ 3 samples out of the total number of landmarks), and the prediction landmark recognition accuracy (the percentage of landmarks with recognition errors ≤ 5 samples out of the total number of landmarks). Extensive simulation tests are conducted on the erroneous data to continuously optimize the project architecture and model features, ultimately resulting in a more complete behavior pattern recognition algorithm and landmark recognition algorithm, thereby improving the accuracy of behavior recognition results.

[0018] In conjunction with the first aspect, in one possible implementation, the behavioral semantic feature data further includes at least one of a third feature data, a fourth feature data, and a fifth feature data, wherein: the third feature data characterizes the change in geomagnetic intensity in the environment detected by the electronic device; the fourth feature data characterizes the change in the magnitude of acceleration of the electronic device in the direction of gravity; and the fifth feature data characterizes the change in the number of WIFI networks detected by the electronic device.

[0019] Research has found that when a user moves between floors using an escalator (in this application, escalator refers to an electric escalator), the Earth's magnetic field generally exhibits a generally continuous upward or downward curve. Furthermore, when the elevator doors open or close, significant geomagnetic fluctuations occur within the elevator. Additionally, when a person moves between floors, their acceleration changes in the vertical direction. Since the number of Wi-Fi networks detected by Wi-Fi sensors in electronic devices within the elevator increases or decreases instantaneously at the moment the elevator doors open or close, the third, fourth, and fifth data can also help identify the user's specific behavioral patterns. Moreover, when the behavioral pattern recognition result is a multi-classification result, compared to extracting only the first and second feature data, further extracting all or part of the semantic features from the third, fourth, and fifth data can help electronic devices better obtain the data characteristics corresponding to the user's current behavioral pattern, thereby improving the accuracy of the behavior recognition results.

[0020] In this embodiment, among the first, second, third, fourth, and fifth feature data included in the behavioral semantic features, the electronic device can simultaneously extract all the data features from these data features, or it can extract only some of the data features from these data features. For example, the electronic device can choose not to extract the first and second feature data, but only extract the third feature data. Compared to extracting only the statistical feature data and temporal feature data, the behavior recognition result obtained by combining the third feature data can also achieve a certain degree of accuracy.

[0021] In conjunction with the first aspect, in one possible implementation, the first feature data includes at least one of small-amplitude weightlessness / weight loss ratio, large-amplitude weightlessness / weight loss ratio, and continuous small-amplitude weightlessness / weight loss ratio. The small-amplitude weightlessness / weight loss ratio represents the proportion of the number of times in the first data segment when the electronic device is in a weightlessness / weight loss state and the weightlessness / weight loss value is less than a second threshold, within the total number of times in the first data segment. The large-amplitude weightlessness / weight loss ratio represents the proportion of the number of times in the first data segment when the electronic device is in a weightlessness / weight loss state and the weightlessness / weight loss value is greater than a third threshold, within the total number of times in the first data segment. The continuous small-amplitude weightlessness / weight loss ratio represents the proportion of the number of times in the first data segment corresponding to the maximum duration of continuous weightlessness / weight loss of the electronic device within the first data segment, within the total number of times in the first data segment. The first data segment is any data segment in the target data.

[0022] As explained above, the human body generally does not experience more than one consecutive data frame of weightlessness or weightlessness when walking. However, the human body can experience momentary and relatively large-scale weightlessness and weightlessness. When the human body is in an elevator, although the human body is mostly stationary relative to the elevator, the human body will also experience a motion state of slight weightlessness (weightlessness) – constant speed – slight weightlessness (weightlessness) as the elevator moves along with it. The state of weightlessness and weightlessness generally lasts for about 2-4 seconds.

[0023] In this embodiment, the difference between the percentage of small-amplitude weight gain / loss and the percentage of continuous small-amplitude weight gain / loss is as follows: the percentage of small-amplitude weight gain / loss considers whether the weight gain / loss of the electronic device at each moment within a certain period is within the small-amplitude weight gain / loss range, without considering the range of the previous moment and the next moment. After obtaining the number of all moments within the small-amplitude weight gain / loss range within a short period, the ratio obtained by dividing the value by the total number of moments within that period is the percentage of small-amplitude weight gain / loss. On the other hand, the percentage of continuous small-amplitude weight gain / loss focuses on the longest consecutive number of moments within the small-amplitude weight gain / loss range within that period, and then the ratio obtained by dividing the number of moments by the total number of moments within that period is the percentage of continuous small-amplitude weight gain / loss.

[0024] In this embodiment, by extracting one or more behavioral semantic features from the proportions of small-amplitude weight gain / loss, large-amplitude weight gain / loss, and continuous small-amplitude weight gain / loss, the electronic device can more accurately identify the behavioral patterns of going up and down elevators and walking on level floors.

[0025] In conjunction with the first aspect, in one possible implementation, the second feature data includes at least one of single-sample maximum turning and double-sample maximum turning, wherein the single-sample maximum turning represents the user's maximum turning angle during the sampling duration corresponding to a data frame; and the double-sample maximum turning represents the maximum value of the sum of the turning angles of any two consecutive sampling moments of the user in a data frame.

[0026] When a person moves from one floor of a staircase to an adjacent floor, they need to make a 180° turn. When climbing multiple floors, this requires multiple 180° turns. Furthermore, the time interval between each consecutive 180° turn is generally short. Considering common behavior, after entering an elevator, a person typically turns to face the elevator door and presses the button for their desired floor, and then likely remains facing the door until exiting. Once at their floor and out of the elevator, they will generally walk straight, turn left, or turn right, but regardless of their direction after exiting the elevator, they generally will not make another 180° turn. Even if a 180° turn occurs again in subsequent actions, the time interval between this turn and the previous one may be relatively long.

[0027] Therefore, in this embodiment, by extracting one or more behavioral semantic features from the single-sample maximum turning and the double-sample maximum turning, the electronic device can more accurately identify the behavior patterns of going up and down stairs.

[0028] Specifically, the maximum turning of a single sample satisfies:

[0029] MaxAngle = Max(Angle);

[0030]

[0031] The dual-sample maximum turning condition satisfies:

[0032] MaxDoubleAngle=Max(Angle (t-1) +Angle t );

[0033] Wherein, abs represents the absolute value function, degress represents the function used to convert radian values ​​into corresponding angles, and gyr x The gyr y The gyr z These represent the angular velocities of the gyroscope sensor rotating around the X-axis, the Y-axis, and the Z-axis, respectively; the acc x The aforementioned acc y The aforementioned acc z These represent the components of the accelerometer on the X-axis, Y-axis, and Z-axis, respectively, and max represents the function for finding the maximum value.

[0034] In conjunction with the first aspect, in one possible implementation, the third feature data includes a mean geomagnetic fluctuation, which represents the average value of the degree of data variation corresponding to a plurality of geomagnetic data segments of second length. The degree of data variation corresponding to any geomagnetic data segment among the plurality of geomagnetic data segments of second length is determined by the maximum and minimum geomagnetic values ​​in that geomagnetic data segment. The plurality of geomagnetic data segments of second length are all included in a first data segment, which is any data segment in the target data.

[0035] Studies have found that when users move between floors using escalators (in this application, escalator refers to electric escalator), the Earth's magnetic field generally exhibits a generally continuous upward or downward curve; in addition, when the elevator doors open or close, intense geomagnetic fluctuations are generated in the elevator.

[0036] Therefore, in this embodiment, by extracting the average value of the geomagnetic fluctuations, the electronic device can more accurately detect the behavior patterns of going up and down stairs.

[0037] Specifically, the mean value of the geomagnetic fluctuations satisfies:

[0038]

[0039] windowSize = winSize * 100;

[0040]

[0041] MagSlice=Mag[t:t+windowSize];

[0042] Wherein, winsize represents the window size corresponding to the first data segment, the first data segment is any data segment in the target data obtained based on the data collected by the geomagnetic sensor, T represents the total amount of geomagnetic data in the first data segment; MagSlice represents all data in the first data segment from time t to (t+windowSize); Max and Min represent the maximum and minimum values ​​of the geomagnetic data collected in the first data segment from time t to (t+windowSize), respectively; and MagDiffAvg is the mean geomagnetic fluctuation corresponding to the first data segment.

[0043] In conjunction with the first aspect, in one possible implementation, the fifth feature data includes a WIFI change ratio, which represents the ratio of the number of first WIFI networks to the number of second WIFI networks, wherein the number of second WIFI networks is the number of WIFI network addresses contained in the last data frame of the first data segment, and the number of first WIFI networks is the number of WIFI network addresses that exist in the last data frame of the first data segment but do not exist in the first data frame of the first data segment.

[0044] The WIFI change ratio feature reflects the change in the number of WIFI networks detected by electronic devices. Since the number of WIFI networks detected by the WIFI sensors in electronic devices inside the elevator will increase or decrease instantly at the moment the elevator door closes or opens, in this embodiment, the behavioral semantic feature of the WIFI change ratio can be well used to identify whether a user is taking the elevator.

[0045] In a second aspect, embodiments of this application provide an electronic device, the electronic device comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method in the first aspect or any possible implementation of the first aspect.

[0046] Thirdly, embodiments of this application provide a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform a method as described in any possible implementation of the first aspect.

[0047] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform a method as described in any possible implementation of the first aspect. Attached Figure Description

[0048] Figure 1 A schematic diagram illustrating a process for identifying a user's behavioral state, provided as an embodiment of this application;

[0049] Figure 2 A schematic diagram illustrating the training / use process of a prediction model provided in an embodiment of this application;

[0050] Figure 3 A schematic diagram illustrating a data segmentation scenario provided in an embodiment of this application;

[0051] Figure 4 A schematic diagram illustrating data collected by an accelerometer according to an embodiment of this application;

[0052] Figure 5 A schematic diagram illustrating the acceleration variation trend of an electronic device under various motion states, provided as an embodiment of this application;

[0053] Figure 6 This application provides a schematic diagram of a human climbing stairs as an embodiment of the present application.

[0054] Figure 7 This application provides a schematic diagram of a human riding an elevator, as illustrated in an embodiment of the present application.

[0055] Figure 8 A schematic diagram illustrating the changing trends of geomagnetic values ​​under various motion states, provided as an embodiment of this application;

[0056] Figure 9 A logic diagram illustrating the determination of the WIFI change ratio provided in an embodiment of this application;

[0057] Figure 10 A schematic diagram illustrating the relationship between sample window length and model accuracy, provided for an embodiment of this application;

[0058] Figure 11 This is a schematic diagram illustrating a scenario for data segmentation of first data, provided as an embodiment of this application.

[0059] Figure 12 A schematic diagram of real landmark data and predicted landmark data provided in an embodiment of this application;

[0060] Figure 13 An architecture diagram of a behavior recognition algorithm provided in an embodiment of this application;

[0061] Figure 14 An architectural diagram of an electronic device provided in an embodiment of this application;

[0062] Figure 15 A flowchart illustrating an identification method provided in an embodiment of this application;

[0063] Figure 16 This is a flowchart of a data processing method provided in an embodiment of this application. Detailed Implementation

[0064] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.

[0065] This application relates to identification methods and electronic devices. For ease of understanding, the relevant terms involved in the embodiments of this application will be introduced below.

[0066] (1) Behavior recognition and behavior recognition results

[0067] Behavior recognition aims to analyze and identify human action types and behavioral patterns through a series of observations, and then describe them using natural language and other methods. Due to the complexity and diversity of human behavior, the recognition results are often diverse, and include probabilistic outputs of behavior types. Methods for recognizing human motion behavior features can be broadly categorized into two main types: machine vision-based and sensor-based.

[0068] Machine vision-based recognition relies on image processing technology to describe and identify human behavior frame-by-frame in video sequences. Its processing is complex, computationally intensive, and easily affected by environmental factors. Sensor-based recognition, on the other hand, offers advantages such as low cost, small size, and less susceptibility to environmental influences. With the rapid development of information technology, mobile and wearable devices are growing at an accelerated pace, and their performance and embedded sensors are becoming increasingly diverse, including high-definition cameras, light sensors, gyroscopes, accelerometers, GPS, and temperature sensors. These various sensors constantly record user information, which can be used not only for predicting user location but also for recognizing user behavior.

[0069] In this application, the behavior recognition result may include the user's behavior pattern recognition result. Optionally, the behavior recognition result may also include landmark (initial and ending points of elevators and escalators) recognition results.

[0070] The behavior pattern recognition result is the user's current behavior pattern. Specifically, in this embodiment, the user's behavior pattern can be divided into a horizontal behavior pattern and a vertical behavior pattern. The horizontal behavior pattern can include being stationary and walking, while the vertical behavior pattern can include going up and down stairs, going up and down elevators, and going up and down escalators. Optionally, when the prediction model 203 predicts the user's behavior pattern, the prediction model can use a multi-class to binary classification approach. First, it determines which of the five behavior patterns (stationary, walking, going up and down stairs, going up and down elevators, and going up and down escalators) the user's action belongs to. Then, it processes the results according to requirements, merging the recognition results, classifying being stationary and walking as the horizontal behavior pattern, and going up and down stairs, going up and down elevators, and going up and down escalators as the vertical behavior pattern.

[0071] The landmark recognition module outputs landmark point recognition results. These results can reflect the moments when user behavior switched as described in the behavioral pattern recognition results. For example, if a sequence of numbers indicates that a user successively performed a horizontal movement mode and a cross-level movement mode, the landmark point recognition results can indicate at what moment the user ended the horizontal movement and began the cross-level movement.

[0072] (2) Accelerometer

[0073] An accelerometer is a sensor that measures acceleration. It typically consists of a mass, a damper, an elastic element, a sensing element, and adaptive circuitry. During acceleration, the sensor measures the inertial force acting on the mass and uses Newton's second law to obtain the acceleration value. Depending on the sensing element, common accelerometers include capacitive, inductive, strain gauge, piezoresistive, and piezoelectric types.

[0074] Most accelerometers on mobile phones are currently capacitive. Their design principle is as follows: a mass can move along a certain axis, a miniature spring is fixed at one end and connected to the mass at the other end, the mass is fixed to one plate of a capacitor, and the other plate of the capacitor is fixed. When the phone accelerates in that direction, the mass compresses or stretches the spring, causing a change in the distance between the capacitor plates, resulting in a change in capacitance. If the capacitor's charge remains constant, the voltage between the plates will change with the change in distance. In this way, the acceleration signal in a certain direction of the phone is converted into a voltage signal. Modern accelerometers typically integrate sensors for the X, Y, and Z axes (three mutually perpendicular directions) together, enabling the sensor to detect the phone's movement in three-dimensional space.

[0075] (3) Gyroscope sensor

[0076] A gyroscope sensor is a simple and easy-to-use positioning and control system based on free-space movement and gestures. Originally used in helicopter models, it is now widely used in mobile portable devices such as smartphones. A gyroscope, also called an angular velocity sensor, can accurately measure the rotation and deflection of a phone, thus enabling corresponding operations. The principle of a gyroscope is that the direction pointed to by the axis of rotation of a rotating object will not change when not affected by external forces. Based on this principle, it is used to maintain orientation. Various methods are then used to read the direction indicated by the axis, and the data signal is automatically transmitted to the control system. Therefore, gyroscope sensors are commonly used to sense how a device is held and to determine the orientation of moving objects.

[0077] (4) Geomagnetic sensor

[0078] Geomagnetic sensors determine magnetic field strength by measuring changes in resistance and are mostly used in compasses and map navigation. Geomagnetic sensors employ anisotropic magnetoresistance (AMR) materials to detect the magnitude of magnetic field strength in space. This crystalline alloy material is highly sensitive to external magnetic fields; changes in the strength of the magnetic field cause changes in the AMR's resistance. Therefore, when the magnetic field around an electronic device equipped with a geomagnetic sensor changes, even a small change, the sensor can sensitively detect the change in magnetic field strength.

[0079] (5) WIFI sensor

[0080] A Wi-Fi sensor can be used to search for / scan for existing Wi-Fi network signals and can be used to obtain the physical address of each Wi-Fi signal. It should be understood that the term "Wi-Fi sensor" in this application is used only to refer to a component used for searching for / scanning existing Wi-Fi network signals. The Wi-Fi sensor can be a separate processing element or it can be implemented on the same chip as other elements (such as the other sensors described above). Furthermore, the Wi-Fi sensor can also be stored as program code in the controller's storage element, and its functions can be called and executed by a processing element of the processor.

[0081] (6) Temporal characteristics and statistical characteristics

[0082] Movements are the most basic units of motion and can be recognized without relying on other movements that occur before or after them. However, an activity is composed of a continuous sequence of movements.

[0083] The temporal characteristics of behavioral actions represent the changes caused by the accumulation of multiple actions performed by a user over a period of time. These changes are specifically reflected in the user's displacement and posture, and can be quantified by integrating data collected from various sensors. For example, performing single and double integrations on acceleration and gyroscope (angular velocity) data respectively yields results that can be roughly understood as the user's velocity and displacement. Based on different human behaviors, such as the significantly smaller velocity and displacement when the human body is stationary compared to when walking, this helps the model to make a rough classification of human behavior. The statistical characteristics of behavioral actions reflect the interrelationships between the overall data contained in the action sequence, considering the changes in data over the entire time period from an individual perspective to a holistic perspective.

[0084] Therefore, when identifying the behavior of a target object, it is necessary to first obtain data that can characterize the action sequence of the target object, and then determine the temporal relationship of the actions in the action sequence and the statistical relationship of each action based on the data, so as to obtain the temporal and statistical features of the action sequence and thus obtain the behavior action identification result.

[0085] (7) Behavioral semantic features

[0086] Compared to actions and movements, behavior is a broader concept, typically encompassing interaction with the environment and causal relationships. That is, when a person performs a certain action, especially when changing floors, the characteristics of their environment will prompt specific actions, and this behavior will also cause changes in certain parameters. For example, when a user walks up or down stairs, upon reaching the end of a staircase and entering the next, the user may need to turn approximately 180 degrees. Similarly, when a user enters an elevator, if the elevator is moving upwards, its motion follows a pattern of acceleration, reaching maximum speed, then constant speed, deceleration, and finally reaching zero speed. According to Newton's laws, during the acceleration phase, the user will experience a state of weightlessness; conversely, during the deceleration phase, the user will experience a state of weightlessness. These states of weightlessness and weightlessness generally last for 2-3 seconds. Therefore, relying solely on temporal and statistical features to identify human behavior cannot guarantee accurate results. Based on this, electronic devices can combine the differences in details and data of various behaviors, that is, the semantic features of behavioral actions, to identify user behavior actions, thereby improving the accuracy of the identification results.

[0087] According to relevant research, people spend more than 80% of their time indoors each day. Therefore, research on indoor behavior is crucial. Accurately identifying people's indoor behavior can significantly advance many research fields. For example, it can help study whether people switch floors, or even monitor the health of elderly people and children at home, checking for falls or other dangerous actions.

[0088] In today's society, powerful smart devices are becoming increasingly common, gradually becoming indispensable necessities in people's daily lives. People generally carry smart devices with them wherever they are, making them a tool that can discreetly identify people's behavioral habits anytime, anywhere. Over the years, the number of sensors integrated into smart devices has increased with the continuous updates and iterations of mobile phones. How to utilize the device's motion sensors and radio frequency signal information to obtain people's current movement status has become a hot topic in technological research, and behavior recognition, as a branch of technological research, is also constantly developing.

[0089] Currently, the main method for studying human behavior using smart devices is to use data from various sensors in smart devices (or features extracted from these sensor data) to train a machine learning model, and then use the trained machine learning model to predict the current state of human behavior.

[0090] Figure 1This diagram illustrates the process by which smart devices use data collected by sensors to identify the user's behavioral state. For example... Figure 1 As shown, the data acquisition module contains most of the sensors found in smart devices, except for... Figure 1 In addition to the barometric pressure sensor, Wi-Fi sensor, and accelerometer shown, these sensors can also include pressure sensors, gyroscopes, and other similar devices. During behavior state recognition, the smart device preprocesses the raw data collected by these sensors through filtering and noise reduction. The resulting data is then transmitted to the feature engineering module, which analyzes and processes this data, performing operations such as integration and averaging, to obtain temporal and statistical features that characterize user behavior. Furthermore, the smart device converts these temporal and statistical features into digital features suitable for machine learning and inputs them into a pre-trained prediction model, which then outputs the final recognition result.

[0091] However, combined Figure 1 As explained above, current smart devices rely on extensive sensor data to acquire information about the current human behavior state, which undoubtedly increases power consumption and reduces recognition efficiency. Furthermore, current prediction models lack the ability to mine semantic features of behavior during training and use. When some sensors fail or are unavailable, the model's accuracy decreases. For example, when studying whether a person is engaging in cross-level behavior, the absence of barometric pressure sensor data could cause the prediction model to fail, resulting in inaccurate predictions.

[0092] To address the aforementioned shortcomings, this application provides an identification method and an electronic device. The electronic device deploys the prediction model provided in this application, enabling it to identify user actions based on the identification method. During the identification process, the electronic device utilizes behavioral features characterizing user actions as input to the prediction model, obtaining more reliable data with fewer sensors. This improves the accuracy of the identification results while reducing the device's power consumption.

[0093] First, a schematic diagram of the training / application process of the prediction model in the electronic device provided in this application is introduced.

[0094] It should be understood that the biggest difference between the training and application of a prediction model lies in the difference in data. However, the data processing logic of the training and application processes of a prediction model is quite similar. Therefore, the following section will take the application process of the prediction model as an example to introduce the prediction model provided in this application.

[0095] like Figure 2As shown, the data used by the prediction model provided in this application is mainly acquired by sensors in the data acquisition module 201 of the electronic device. Specifically, the sensors in the data acquisition module may include an accelerometer 201a, a gyroscope 201c, a geomagnetic sensor 201b, and a WIFI sensor 201d. Among them, the accelerometer 201a, the gyroscope 201c, and the geomagnetic sensor 201b can also be referred to as a 3-axis accelerometer, a 3-axis gyroscope, and a 3-axis electronic compass (geomagnetic sensor), respectively. The combination of these three sensors can be referred to as a nine-axis sensor.

[0096] It is understood that the structure illustrated for the data acquisition module 201 in this embodiment does not constitute a specific limitation on the data acquisition module 201. In other embodiments of this application, the data acquisition module 201 may include more or fewer sensors than illustrated, some or all of which may be integrated into a single element or implemented in hardware, software, or a combination of both.

[0097] It should be noted beforehand that during the training process of the aforementioned predictive model, the training dataset used can come from multiple devices. That is, the manufacturer of the aforementioned electronic devices can use data sets collected from multiple devices as the training set for the predictive model, which is then trained by devices equipped with model training capabilities. Subsequently, during the production process of the aforementioned electronic devices, the predictive model is uniformly deployed across the devices. Then, when the electronic devices deploying this predictive model recognize user actions (i.e., during the use of the predictive model), each electronic device can use only its own collected sensor data as input to the predictive model to obtain the recognition result of the user's actions.

[0098] Specifically, during the process of electronic devices using predictive models to identify user behavior, each sensor in the data acquisition module can collect data at a certain sampling frequency. This data directly collected by each sensor can be referred to as initial data. Optionally, the sampling frequency of each sensor can be 100Hz, that is, 100 data collections per second.

[0099] After collecting the initial data, the electronic device can perform preprocessing such as filtering and denoising on the initial data, and then select a sliding window of a certain length and a certain window coverage to segment the initial data collected by the sensor. That is, the initial data of the long time series is divided into smaller data frames (the length of each data frame is the length of the sliding window mentioned above; in some embodiments, the length of the data frame can also be called the first length) for processing, so that the feature engineering 202 can extract features from the data.

[0100] It is important to understand that the length of the data frame, i.e., the length of the sliding window, directly affects not only the quality of feature extraction and classification performance, but also whether the method is suitable for real-time systems. For example, if the sliding window is too small, each segmented data frame may only represent a part of the action, failing to reflect the complete behavioral state of the action. This results in the extracted features not effectively representing the action, leading to a decrease in recognition performance, and the electronic device needs to perform frequent recognition, increasing the computational load. Conversely, if the sliding window is too large, a single data frame may contain multiple actions, causing the system to be unable to effectively recognize the current behavior, affecting the system's real-time performance, and also impacting the accuracy of the recognition results. Therefore, when segmenting the initial data, the length of the sliding window directly affects the accuracy of the recognition results.

[0101] In addition, the sliding window coverage is also an important factor affecting the accuracy of the recognition results. The sliding window coverage refers to the overlap rate between two adjacent segmentation windows during the data segmentation process. A coverage rate of 0% means that adjacent windows do not overlap, and a coverage rate of 50% means that the current window contains half of the data from the previous window.

[0102] Current research indicates that a window length of 2.56s is a relatively ideal window length because Fourier transform is required for frequency domain calculations. Using a power of 2 as the time window ensures the integrity of the data involved in the frequency domain calculations. Furthermore, a 50% window coverage effectively reduces disturbances caused by transient behavior. Therefore, in this embodiment, the electronic device can use a sliding window with a window length of 2.56s and a window coverage of 50% to segment the initial data. For each sensor, the acquired data, after segmentation, will yield multiple data segments, one of which can be called a "data frame". It is easy to see that when the electronic device samples at a frequency of 100Hz and uses a sliding window with a window length of 2.56s and a window coverage of 50% to segment the initial data acquired by the sensor, each data frame (which can also be called a sample during model training and testing) contains data from 256 time points.

[0103] by Figure 3 The following diagram illustrates a data segmentation scenario. Figure 3In this context, the initial data 30 can be the initial data obtained by sampling continuously for 2 minutes using one or more sensors (such as an accelerometer, gyroscope, geomagnetic sensor, or WIFI sensor) in an electronic device at a sampling frequency of 100 Hz. Assuming the electronic device uses a sliding window with a window length of 2.56 s and a window coverage of 50% to segment the initial data 30, it is easy to calculate that the initial data 30 can be divided into approximately 93 data frames, each containing data from 256 time points. Figure 3 Data frames 301-304 shown are partial consecutive data frames obtained by segmenting the initial data 30. Data frame 301 corresponds to a sampling time of 0s-2.56s, and data frame 302 corresponds to a sampling time of 1.28s-3.84s. That is, the data collected by each sensor in the electronic device within the 1.28s-2.56s time period is included in both data frame 301 and data frame 302, and the length of this data portion accounts for (1.28 ÷ 2.56) × 100% = 50% in both data frames 301 and 302. Similarly, data frames 303 and 302, 304 and 303, and any two subsequent adjacent data frames all exhibit a 50% data overlap.

[0104] After segmenting the data collected by sensors in the electronic device, the device will further transmit the segmented data to feature engineering 202 for feature extraction. In the embodiments of this application and subsequent embodiments, the data obtained after segmenting the initial data and transmitting it to feature engineering 202 can be referred to as first data. It can be understood that the first data is data in units of data frames.

[0105] Feature extraction refers to the process of calculating and generating feature data that can characterize actions from segmented data frames. In this application, in order to focus on studying human behavior patterns and accurately identify five behaviors of the human body: stillness, walking, climbing stairs, climbing escalators, and climbing elevators, the electronic device extracts not only physically meaningful temporal features and statistically meaningful statistical features from the segmented data, but also behavioral semantic features. These behavioral semantic features may include one or more features reflecting the electronic device's (user's) weightlessness / weightlessness state, the turning amplitude of the electronic device (user), the changes in the geomagnetic intensity of the environment in which the electronic device (user) is located, the changes in the magnitude of the acceleration of the electronic device (user) in the direction of gravity, and the changes in the number of WIFI networks in the environment in which the electronic device (user) is located. In other words, the features extracted from the data collected and segmented by the data module by feature engineering 202 can be roughly divided into three categories: statistical features, temporal features, and behavioral semantic features.

[0106] When the data acquisition module 201 includes four sensors: accelerometer 201a, gyroscope 201c, geomagnetic sensor 201b, and WIFI sensor 201d, the electronic device can refer to the following table 1 for the feature content included in the above three types of features (statistical features, temporal features, and behavioral semantic features):

[0107] Table 1

[0108]

[0109]

[0110] As shown in Table 1, when the data acquisition module 201 includes four sensors—accelerometer 201a, gyroscope 201c, geomagnetic sensor 201b, and WIFI sensor 201d—the statistical features extracted by feature engineering 202 are applied to calculate the statistical results of five data points—horizontal acceleration, vertical acceleration, acceleration vector sum, gyroscope vector sum, and geomagnetic vector sum—across 11 dimensions, resulting in a total of 55 features. These 11 dimensions include mean, standard deviation, variance, median, minimum, maximum, difference between maximum and minimum, quartile, kurtosis, skewness, and root mean square. Statistical features are simple to calculate and require minimal computation, but they reflect the interrelationships between the overall data, helping the prediction model to consider data changes over the entire time period from an individual perspective to a holistic one.

[0111] The temporal features extracted by Feature Engineering 202 include four dimensions: the first integral of the acceleration vector sum, the second integral of the acceleration vector sum, the first integral of the gyroscope (angular velocity) vector sum, and the correlation coefficients between the acceleration vector sum and the Y-axis and Z-axis. These features reflect the user's movement speed and displacement. For example, the speed and displacement of a person at rest are significantly less than those during walking. Furthermore, the acceleration YZ correlation coefficient refers to the correlation coefficient between the acceleration along the Y-axis and Z-axis. When the phone is held at a flat position, the correlation coefficients between the Y-axis and Z-axis differ depending on the periodicity of the human body during walking and climbing stairs, as well as the varying stride lengths. Therefore, these temporal features can help the prediction model make a rough classification of human behavior.

[0112] The behavioral semantic features extracted by feature engineering 202 include one or more of the following 11 dimensions: percentage of slight weight gain, percentage of significant weight gain, percentage of slight weightlessness, percentage of significant weightlessness, percentage of continuous slight weight gain, percentage of continuous slight weightlessness, maximum turning angle in a single sample, maximum turning angle in two samples, mean geomagnetic fluctuation, projection of acceleration onto the direction of gravity, and change in WIFI. Specifically, the percentages of slight weight gain, significant weight gain, slight weightlessness, significant weightlessness, continuous slight weight gain, continuous slight weightlessness, and projection of acceleration onto the direction of gravity can be obtained based on initial data collected by accelerometer 201a; the maximum turning angle in a single sample and the maximum turning angle in two samples can be obtained based on initial data collected by gyroscope sensor 201c; the mean geomagnetic fluctuation can be obtained based on initial data collected by geomagnetic sensor 201b; and the change in WIFI can be obtained based on initial data collected by WIFI sensor 201d.

[0113] These behavioral semantic features reflect the differences in details and data exhibited by various behaviors. In other words, the magnitude and / or form of data changes corresponding to these behavioral semantic features will differ when a user performs different actions; or, when a user performs a certain action, the data value and / or form of data change corresponding to a particular feature of these behavioral semantic features will be significantly different from the data value and / or form of numerical change when the user performs other actions. The following will combine... Figures 4-9 The above 11-dimensional behavioral semantic features are explained in detail.

[0114] 1) Percentage of slight overweight, percentage of significant overweight, percentage of slight weight loss, percentage of significant weight loss, percentage of continuous slight overweight, percentage of continuous slight weight loss

[0115] The percentages of slight overweight, significant overweight, slight weightlessness, significant weightlessness, continuous slight overweight, and continuous slight weightlessness reflect the overweight / weightlessness characteristics of electronic devices (users). They are mainly used to identify the user's walking or elevator behavior patterns.

[0116] Studies have found that the human body generally does not experience more than one consecutive data frame of weightlessness or weightlessness when walking. However, the human body can experience momentary and relatively large-scale weightlessness and weightlessness. When a person is in an elevator, although the person is mostly stationary relative to the elevator, the person will experience acceleration and deceleration as the elevator moves. Therefore, the person will also experience a motion state of slight weightlessness (weightlessness) – constant speed – slight weightlessness (weightlessness) as the elevator moves. The state of weightlessness and weightlessness generally lasts for about 2-4 seconds.

[0117] According to basic physics, if an object is only subject to gravity and a supporting force, then when the acceleration of the object in the direction of gravity is 0, it means that the object is in equilibrium under the direction of gravity, and the object is neither in a state of excessive weight nor weightlessness; when the acceleration of the object in the direction of gravity is greater than 0, it means that the gravity acting on the object is greater than the supporting force, and the object is in a state of weightlessness; when the acceleration in the direction of gravity is less than 0, it means that the gravity acting on the object is less than the supporting force, and the object is in a state of excessive weight. Figure 4 This diagram illustrates the data collected by an acceleration sensor as a user enters the elevator on a level floor, moves to their destination, and exits the elevator. For ease of understanding, Figure 4 In the coordinate system shown, the horizontal axis represents time, and the vertical axis represents the user's acceleration in the direction of gravity (or the acceleration of the electronic device in the direction of gravity as measured by the accelerometer). The times t41-t42 and t45-t46 represent the time the user spends walking on the floor, while t42-t45 represents the time the user spends moving with the elevator. In other words, t42 is the time the user enters the elevator, and t45 is the time the user exits the elevator. Figure 4 The images corresponding to times t41-t42 and t45-t46 show that when a user walks on a level surface, they experience momentary and significant weightlessness and g-forces. These phenomena are instantaneous and typically end within one second. However, from... Figure 4 The graphs corresponding to times t42-t45 show that when the user moves with the elevator, they experience slight weightlessness and g-forces. The magnitude of this weightlessness is significantly smaller than that experienced when walking on level ground, and this weightlessness typically lasts for a longer period. Specifically, in... Figure 4Between t44 and t45, the user is in a state of weightlessness, and the elevator may be accelerating downwards at this time; Figure 4 During the period from t42 to t43, the user is in an overloaded state, and the elevator may be slowing down downwards at this time.

[0118] Figure 5 Schematic diagrams illustrating the acceleration variation trends of electronic devices under various motion states are provided.

[0119] Figure 5 Figure (A) shows a schematic diagram illustrating the trend of acceleration changes measured by electronic devices when a user rides an elevator. It should be noted that, for ease of understanding, Figure 5 The graph (A) shown in the figure represents the trend of acceleration change under ideal conditions (i.e., the electronic device is only subjected to gravity and support force), compared to... Figure 4 The image corresponding to the time interval t42-t45. Figure 5 The image shown in (A) is clearer and more intuitive. Combined with the foregoing explanation, it can be seen that... Figure 5 In (A), times t51 and t54 represent the moments when the elevator begins and stops moving, respectively. During the period from t51 to t52, the user accelerates downwards with the elevator, experiencing weightlessness. Conversely, during t53 to t54, the user decelerates downwards with the elevator, experiencing weightlessness. It's important to understand that the time intervals between t51 and t52, and between t53 and t54, are typically quite long, usually 2-3 seconds, and the maximum acceleration the user can achieve during the elevator's movement is m (in the same direction as gravity) and n (opposite to gravity).

[0120] Figure 5 Figure (B) shows a schematic diagram illustrating the acceleration variation trend measured by the electronic device when the user walks on a level surface. As explained above, the user's (or the electronic device's) acceleration fluctuates significantly during this action; that is, there are momentary and substantial instances of weightlessness and hypergravity, but these periods generally do not exceed one consecutive data frame. Furthermore, the maximum acceleration achievable by the user while walking on a level surface is M (in the same direction as gravity) and N (opposite to gravity), and generally, M is greater than m, and N is greater than n.

[0121] Figure 5 (C) in the diagram illustrates the trend of acceleration measured by the electronic device when the user is stationary on a level floor. It is easy to understand that when stationary on a level floor, the user (or electronic device) is in equilibrium under the direction of gravity, and the object is neither in a state of excessive weight nor weightlessness. Therefore, the acceleration of the electronic device is 0 at this time.

[0122] Therefore, based on the above explanation, it can be seen that extracting the proportion of weightlessness at different levels can accurately identify whether a person is in an elevator or is moving or stationary on a floor. For example, if the extracted behavioral semantic features reflect that the electronic device is in a state of weightlessness / weightlessness for a duration of 2-3 seconds, then the electronic device (or prediction model 203) can identify that the user may be taking an elevator and moving between floors.

[0123] It should be noted that the difference between the percentage of slight overweight (weightlessness) and the percentage of continuous slight overweight (weightlessness) is as follows: The percentage of slight overweight (weightlessness) considers whether the electronic device's overweight or weightlessness range is within a slight overweight or weightlessness range at each moment within a certain period of time. It does not need to consider the range of the previous moment and the next moment. After obtaining the number of all moments within the slight overweight or weightlessness range within a short period of time, the ratio obtained by dividing the number of moments within the short period of time is the percentage of slight overweight (weightlessness). On the other hand, the percentage of continuous slight overweight or weightlessness focuses on the longest consecutive number of moments within the slight overweight or weightlessness range within a certain period of time. The ratio obtained by dividing the number of moments within the short period of time is the percentage of continuous slight overweight or weightlessness.

[0124] 2) Maximum turning of single sample and maximum turning of two samples

[0125] Single-sample maximum turning feature and two-sample maximum turning feature are features that reflect the magnitude of changes in the direction of electronic devices (users), and they can be mainly used to identify human behavior of going up and down stairs.

[0126] Figure 6 This is a schematic diagram illustrating a human climbing stairs, as provided in this application. Figure 6 As shown in (A), the human body is currently walking upwards and has reached the top of a certain staircase. If the destination has not yet been reached, the human body needs to climb the next staircase to reach a higher floor. Considering the current structure of staircases, most staircases have a 180-degree turn between adjacent floors. Therefore, when a human body moves from one staircase to an adjacent one, the human body needs to make a 180-degree turn. For example... Figure 6 As shown in (B) in the image, the human body has now transitioned from... Figure 6 In step (A), the person turns 180° in the direction they are facing to enter the next floor. Similarly, after reaching the top of this floor, to reach a higher floor, the person needs to turn 180° again. Furthermore, it should be noted that we assume... Figure 6 The moment when the human body, as shown in (A), undergoes a 180-degree turn is t61, while Figure 6As shown in (B), the moment when the human body makes a 180-degree turn is t62. Generally speaking, the time interval between t61 and t62 is not too long, usually 5-10 seconds. That is to say, when the human body moves across floors while walking on stairs, the time interval between each two adjacent 180-degree turns is generally small.

[0127] Figure 7 This is a schematic diagram illustrating a scenario of a person riding an elevator, as provided in this application. Based on common behavioral habits, generally speaking, after entering an elevator, a person will turn to face the elevator door and press the button for their target floor, and will likely remain facing the elevator door until they exit the elevator. For example... Figure 7 As shown in (A), the human body has just entered the elevator and is not facing the elevator door at this moment. The body then turns 180 degrees, as... Figure 7 As shown in (B), the person is now facing the elevator door. During the elevator's journey, the person will remain facing the elevator door. Afterwards, as... Figure 7 (C) and Figure 7 As shown in (D), after a person reaches the target floor and exits the elevator, they will generally walk straight, turn left, or turn right. However, regardless of the direction of movement after exiting the elevator, the person will generally not make another 180-degree turn. Even if the person makes another 180-degree turn in subsequent actions, the time between this turn and the previous turn (i.e., Figure 7 The time interval between (A) shown in the figure and the moment when the elevator turns after the user enters the elevator may be long.

[0128] Therefore, based on the differences in the frequency and number of turning actions among different human behaviors, the feature engineering provided in this application extracts single-sample maximum turning features and two-sample maximum turning features. These two behavioral semantic features can be used as key features for analysis when users walk up / down stairs or up / down elevators. For example, when the extracted behavioral semantic features reflect that the electronic device performs multiple large-scale turning actions in a short period of time, the electronic device can recognize the user's action pattern to go up / down stairs.

[0129] Specifically, the formula for calculating the maximum turning rate for a single sample is as follows:

[0130]

[0131] MaxAngle = Max(Angle) (1-2)

[0132] Where abs represents the function used to calculate the absolute value, degress represents the function used to convert radian values ​​to their corresponding angles, and gyr x gyr y gyrz These represent the angular velocities of the gyroscope sensor rotating around the X-axis, Y-axis, and Z-axis, respectively; acc x acc y acc z These represent the components of the accelerometer along the X-axis, Y-axis, and Z-axis, respectively.

[0133] Furthermore, when the curvature of a staircase corner is long or the user (e.g., an elderly person or a child) moves slowly, the maximum turning angle achievable in a single sample may not be particularly large. In such cases, electronic devices are highly likely to misinterpret the user's behavior as stationary or moving horizontally. This application can also extract the maximum turning angle from two samples from the data collected by the gyroscope sensor, avoiding incorrect identification results from electronic devices for staircases with large corners or when some users turn slowly. Specifically, the formula for calculating the maximum turning angle from two samples is as follows:

[0134] MaxDoubleAngle = max(Angle) (t-1) +Angle t (1-3)

[0135] Among them, Angle (t-1) Angle t This represents the maximum single-sample turning value corresponding to two adjacent data frames, and max represents the maximum value among the maximum single-sample turning values ​​corresponding to two adjacent data frames.

[0136] 3) Mean value of geomagnetic fluctuations

[0137] The mean geomagnetic fluctuation is a characteristic that reflects the magnitude of geomagnetic changes in the environment in which electronic devices (users) are located. It is mainly used to identify the behavioral patterns of users when taking elevators or escalators.

[0138] Studies have found that when users move between floors using escalators (in this application, escalator refers to electric escalator), the Earth's magnetic field generally exhibits a generally continuous upward or downward curve; in addition, when the elevator doors open or close, intense geomagnetic fluctuations are generated in the elevator.

[0139] Figure 8 A schematic diagram showing the changing trends of geomagnetic values ​​measured by electronic devices under various motion states is presented.

[0140] Figure 8 Figure (A) shows a schematic diagram of the geomagnetic field variation trend collected by electronic equipment when a user moves between floors in an elevator. Figure 8In (A), time t81-t84 represents the period when the user enters the elevator, and time t1-t2 represents the period when the user exits the elevator. During both periods, the elevator door will open and then close. Figure 8 As shown in image (A), the geomagnetic field, which was originally stable, experienced violent fluctuations during the time intervals of t81-t82 and t83-t84. However, during other time intervals, the geomagnetic values ​​measured by the electronic device showed a stable and unchanging trend.

[0141] Figure 8 Figure (B) shows a schematic diagram of the geomagnetic field variation trend collected by electronic equipment when a user moves between floors on an escalator. Figure 8 In (B), time t85-t86 represents the period when the user approaches and steps onto the escalator, and time t86-t87 represents the period when the user moves upwards with the escalator. From Figure 8 As shown in image (B), the originally stable geomagnetic field experienced violent fluctuations during the period from t85 to t86; and during the period from t86 to t87, when the user moved with the escalator, the geomagnetic value showed a generally continuously rising curve.

[0142] Figure 8 Figure (C) shows a schematic diagram of the geomagnetic field changes collected by an electronic device when a user moves between floors or goes up and down stairs. During this process, the overall value of the geomagnetic field does not fluctuate.

[0143] Therefore, fluctuations in the Earth's magnetic field can reflect human behavior to some extent. For example, if the extracted behavioral semantic features indicate that the geomagnetic values ​​detected by the electronic device generally show an upward or downward trend over a period of time, the electronic device can identify the user's action patterns to indicate whether they are going up or down an escalator.

[0144] Specifically, the formula for calculating the mean of geomagnetic fluctuations can be shown below:

[0145] windowSize=winSize*100 (1-4)

[0146]

[0147] MagSlice=Mag[t:t+windowSize] (1-6)

[0148]

[0149] Here, Winsize represents the window size used to divide samples into data frames during feature extraction in Feature Engineering 202, i.e., the number of data frames contained in each data set during feature extraction. For example, the specific value of WinSize can be 5, 10, or 15, which will be explained later. T represents the total amount of geomagnetic data, and its specific value is equal to (Winsize * the number of geomagnetic data points contained in each data frame). MagSlice represents the data in the geomagnetic data from time slice t to (t + windowSize). Max and Min represent the maximum and minimum values ​​of the geomagnetic data in the geomagnetic data from time slice t to (t + windowSize), respectively. MagDiffAvg is the mean of the geomagnetic fluctuations.

[0150] 4) Projection of acceleration in the direction of gravity

[0151] When the human body moves across layers, its acceleration changes in the vertical direction. However, the position and orientation of electronic devices are uncertain when a person uses or carries them. Therefore, during feature extraction, the projection of the acceleration onto the Z-axis in the mobile phone's coordinate system cannot be used. However, based on basic physics, acceleration is a vector, and acceleration in one direction can be decomposed into multiple accelerations in other directions. Therefore, during feature extraction, the electronic device can calculate the projection of the acceleration onto the direction of gravity to extract features. Specifically, the electronic device can obtain gravitational acceleration by calculating the average acceleration within a data frame. This is equivalent to performing a simple low-pass filter on the data by calculating the average of a small window of data, which can cancel out the noise caused by the up-and-down swaying of the human body while walking. After obtaining the gravitational acceleration, the electronic device can use vector projection to obtain the projection of the acceleration onto the direction of gravity.

[0152] 5) WIFI change ratio

[0153] The WIFI change ratio is a feature that reflects the magnitude of change in the number of WIFI networks in the environment in which an electronic device is located. It is mainly used to identify the user's behavior patterns when going up / down elevators.

[0154] The WIFI change ratio feature reflects the change in the number of WIFI networks detected by electronic devices. Since the number of WIFI networks detected by the WIFI sensors in electronic devices inside the elevator will increase or decrease instantly at the moment the elevator door closes or opens, this WIFI change ratio feature can be well used to identify whether a user is taking the elevator.

[0155] In this embodiment, the WIFI change ratio is equal to the ratio of the number of newly appearing WIFI networks at the last moment in the WIFI data collected by the WIFI sensor to the total number of WIFI networks at the last moment. The number of newly appearing WIFI networks at the last moment is obtained by comparing it with the number of WIFI networks at the first moment. Here, "first moment" and "last moment" correspond to the first and last data frames in the sample window, respectively. The "sample window" represents the window size used by feature engineering 202 to divide the first data into units of data frames during feature extraction; that is, the number of data frames contained in each group of data during feature extraction. For details, please refer to the subsequent description, which will not be elaborated here. Specifically, the electronic device can first obtain the intersection of the WIFI hardware address at the last moment and the WIFI hardware address at the first moment, and then use the difference between the WIFI hardware address at the last moment and the intersection to obtain the number of newly appearing WIFI networks. Afterwards, the electronic device can divide the number of newly appearing WIFI networks by the total number of WIFI networks at the last moment; the resulting ratio is the WIFI change ratio.

[0156] by Figure 9 Let's take an example to illustrate. Figure 9 In the table, address list 901 contains all the network addresses in the first data frame collected by the electronic device within a certain sample window, i.e. Figure 9 The network addresses shown are MAC1 to MAC7; address list 902 contains all the MAC addresses in the last data frame collected by the electronic device within the above sample window, i.e. Figure 9 The network addresses shown are mac2, mac4, and mac7-mac13; any network address in address list 901 and address list 902 corresponds to a Wi-Fi network, and the same network address in the two address lists corresponds to the same Wi-Fi network.

[0157] As explained above, after the electronic device obtains address lists 901 and 902, it can compare and analyze the network addresses in these two lists to obtain the set of MAC addresses that exist in both lists, namely network address mac2, network address mac4, and network address mac7. Figure 9 The intersection list is 903.

[0158] Then, the electronic device uses the difference between address list 902 and the intersection list to obtain... Figure 9The difference list 904 shows the newly appearing Wi-Fi networks corresponding to the network addresses in the difference list. From difference list 904, we can see that all the newly appearing Wi-Fi networks are those corresponding to network addresses MAC8-MAC13, totaling 6. From address list 902, we can see that the electronic device detected 9 Wi-Fi networks at the last moment. Therefore, combining the above explanations, the Wi-Fi change ratio for this sample window is 6 / 9, or 2 / 3.

[0159] Furthermore, it should be noted that this is because samples extracted from a single data frame cannot accurately reflect the semantic features before and after the current moment. For example, during elevator operation, there are periods of constant speed. If only a single data frame is used to extract features at this moment, it will be found that the features at this moment are not significantly different from those at a stationary plane. Therefore, if only a single data frame is used to extract features at this moment, the information captured by the data within the window is too limited, and the recognition result output by the prediction model may have low accuracy. Thus, in some embodiments, during feature extraction in feature engineering 202, the electronic device uses a sample window based on data frames to segment the aforementioned first data. That is, during feature extraction, the electronic device performs feature extraction in batches of several data frames.

[0160] Furthermore, the duration of the relevant feature data varies for different behavior patterns. For example, when a user is moving between floors in an elevator, the sample window needs to be longer to fully extract feature data indicating whether the user is in a state of weightlessness or weightlessness. However, when the user is in a stationary or moving mode on a level floor, only a few data frames are needed to reflect the user's motion characteristics, so the sample window can be set shorter accordingly. Therefore, in some embodiments, the electronic device can simultaneously set multiple sample windows of different sizes to segment the first data.

[0161] It should be understood that when determining the size of the sample window, the amount of information captured within the sample window needs to be sufficient, but the sample window should not contain too much data of different behaviors. Therefore, in some embodiments, during feature extraction, the electronic device can also set the sample window to different sizes (i.e., there are multiple sample windows, each containing a different number of data frames), and simultaneously use these sample windows of different sizes to segment the same first data, resulting in multiple sets of data (in some embodiments, these multiple sets of data can also be referred to as at least two sets of data segments, and any one of the data segments in these multiple sets of data can be referred to as the first data segment). When performing feature extraction, extracting features from the aforementioned multiple sets of data separately can effectively improve the performance of the prediction model and the accuracy of the recognition results.

[0162] It is important to note that, unlike the segmentation of initial data collected by various sensors using a sliding window, the size of the sample window used for segmenting the first data is determined by the data frame as the smallest unit, and the electronic device can simultaneously use multiple sample windows of different sizes to segment the first data. However, when segmenting the initial data, the window size is determined by the number of time points as the smallest unit, and the coverage between windows can be set to 50%. For example, if the electronic device samples the initial data at a frequency of 100Hz, and the window size determined by the electronic device for the initial data is 2.56s, then when segmenting the initial data, the number of data points in each window is 256. These 256 data points correspond to 256 sampling times. Each data segment obtained from segmenting the initial data is a data frame, and all data frames obtained from the initial data constitute the first data. For details, please refer to the aforementioned... Figure 3 The relevant explanations will not be repeated here. When segmenting the aforementioned first data subsequently, the electronic device can determine the sample window for the first data in units of data frames.

[0163] For example, when segmenting the first data mentioned above, the specific size of the sample window can be set to 5, 10, and 15, respectively; that is, the size of the sample window can be the length corresponding to 5 data frames, 10 data frames, and 15 data frames. These three window sizes were determined through observation of a large amount of experimental data. Figure 10 As shown, when the number of sample windows is 3, the accuracy of the prediction model is the worst when the sample window size is set to 5, 8, or 11, with an accuracy of only about 0.972. However, when the sample window size is set to 5, 10, or 15, the accuracy of the prediction model is the best, with an accuracy of about 0.98.

[0164] Figure 11 The diagram illustrates how an electronic device segments the first set of data using sample windows of sizes 5, 10, and 15. It is assumed that the sensor in the electronic device continuously collects data for 2 minutes, obtaining an initial data segment. The electronic device then segments the initial data using windows of size 2.56 seconds and a window coverage of 50%, ultimately obtaining the following data... Figure 11 The first data 011 shown in (A) above. Simple mathematical calculations show that the first data 011 contains approximately 93 data frames, and the sampling duration of the data contained in each data frame is 2.56s.

[0165] Subsequently, during feature extraction, the electronic device segmented the aforementioned first data using sample windows of sizes 5, 10, and 15, respectively. For example... Figure 11 As shown in (A), when the first data 011 is segmented using a sample window of size 5, the first data segment corresponds to the first to fifth data frames, and the sampling duration of this data segment is 7.68s; when the first data 011 is segmented using a sample window of size 10, the first data segment corresponds to the first to tenth data frames, and the sampling duration of this data segment is 14.08s; when the first data 011 is segmented using a sample window of size 15, the first data segment corresponds to the first to fifteenth data frames, and the sampling duration of this data segment is 20.48s.

[0166] When segmenting the first data using a sample window of a fixed size, the sliding step of the sample window can be 1. Here, we take a sample window size of 5 as an example for illustration. Figure 11 As shown in (B), when the first data 011 is segmented using a sample window with a window size of 5, data segment 1101 is the first data segment obtained by the electronic device from the first data 011, corresponding to the 1st to 5th data frames in the first data 011; data segment 1102 is the second data segment obtained by the electronic device from the first data 011, corresponding to the 2nd to 6th data frames in the first data 011; data segment 1103 is the third data segment obtained by the electronic device from the first data 011, corresponding to the 3rd to 7th data frames in the first data 011, and so on, until the last data segment corresponding to the 89th to 93rd data frames in the first data 011 is obtained. Figure 11 (Not shown in the image).

[0167] Similarly, the specific process of segmenting the first data using sample windows of other lengths can also be referenced from the previous section. Figure 11 The relevant explanations for (B) in the text will not be repeated here.

[0168] After feature engineering 202 completes feature extraction of the target data, the electronic device inputs the temporal features, statistical features, and behavioral semantic features extracted by feature engineering 202 into the prediction model 203. The prediction model 203 analyzes these features and finally outputs the recognition result of the user's behavior. Specifically, the prediction model 203 can use the XGBoost model.

[0169] In this embodiment of the application, the prediction model 203 may include two modules, namely Figure 2 The behavior recognition module and landmark recognition module shown in the diagram, and the behavior recognition results output by the prediction model 203, can also include user behavior pattern recognition results and landmark (initial and ending points of elevators and escalators) recognition results. Wherein:

[0170] The behavior recognition module is used to output the user's behavior pattern recognition result. In some embodiments, this behavior pattern recognition result can be referred to as the first recognition result. The behavior pattern recognition result is the user's behavior pattern at this time. Specifically, in the embodiments of this application, the user's behavior pattern can be divided into a flat-level behavior pattern and a cross-level behavior pattern. Among them, the flat-level behavior pattern can include being stationary and walking, and the cross-level behavior pattern can include going up and down stairs, going up and down elevators, and going up and down escalators. Optionally, when the prediction model 203 predicts the user's behavior pattern, the prediction model can adopt a multi-class to binary classification method, first determining which of the five behavior patterns (stationary, walking, going up and down stairs, going up and down elevators, and going up and down escalators) the user's behavior action belongs to, and then processing according to the requirements, merging the recognition results, that is, classifying being stationary and walking as the flat-level behavior pattern, and classifying going up and down stairs, going up and down elevators, and going up and down escalators as the cross-level behavior pattern.

[0171] Furthermore, the behavior pattern recognition results output by the prediction model 203 can be represented by characters, specifically as a sequence of numbers. In this sequence, each number represents a duration during which the user acted in the corresponding behavior pattern. This duration can be the duration of a data frame. For example, when the behavior pattern recognition result is a multi-class result (i.e., it is necessary to specifically identify which behavior pattern the user is in: stationary, walking, climbing stairs, using an elevator, or using an escalator), the recognition results output by the electronic device can be represented by 0, 1, 2, 3, 4, and 5, respectively. Similarly, when the behavior pattern recognition result is a binary result (i.e., it is only necessary to identify whether the user is in a horizontal or horizontal movement pattern), the recognition results output by the electronic device can be represented by 0 and 1, respectively, to indicate the horizontal and horizontal movement patterns. For example, when the behavior pattern recognition result is a binary classification result, assuming that the behavior pattern recognition result output by the prediction model 203 is a sequence of numbers such as [0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 0, 0], then this sequence of numbers represents the user moving in a flat motion mode for a period of time, which is equal to the duration corresponding to 5 data frames; then, the user moves in a cross-layer motion mode for another period of time, which is equal to the duration corresponding to 6 data frames; finally, the user moves in a flat motion mode for another period of time, which is equal to the duration corresponding to 2 data frames.

[0172] The landmark recognition module outputs landmark point recognition results, which in some embodiments may be referred to as second recognition results. The landmark point recognition results can reflect the moment when user behavior switched in the aforementioned behavior pattern recognition results. For example, if a sequence of numbers reflects that a user successively performed a horizontal movement mode and a cross-level movement mode, the landmark point recognition results can reflect at what moment the user ended the horizontal movement and began the cross-level movement.

[0173] Optionally, once the prediction model 203 has been trained and deployed in an electronic device, when the electronic device subsequently uses the prediction model 203 to perform actions on the user, the output of the prediction model may only include the user's behavior pattern recognition results, without including the landmark point recognition results.

[0174] It should be understood that the landmark recognition result output by the prediction model 203 is the result of the electronic device predicting the user's behavior based on the data collected by the sensor, but there may still be some time error between it and the user's actual landmark point.

[0175] like Figure 12As shown, sequence 012 represents the landmark data that characterizes the user's actual behavior pattern, while sequence 012' represents the landmark data output by the prediction model when predicting the user's behavior pattern. In these two sequences, the number "0" indicates that the user is in a flat-level movement pattern, while the number "1" indicates that the user is in a cross-level movement pattern.

[0176] Sequence 012 shows that there are three time periods corresponding to the user's actual cross-layer movement, which are the data segments S1-E1, S2-E2, and S3-E3 in sequence 012. Similarly, sequence 012' shows that the predictive model also identifies three time periods corresponding to cross-layer movement in user behavior, which are the data segments S1'-E1', S2'-E2', and S3'-E3' in sequence 012. Specifically, the S1-E1, S2-E2, and S3-E3 segments in sequence 012 correspond to the S1'-E1', S2'-E2', and S3'-E3' segments in sequence 012', respectively. Understandably, in the data segments corresponding to sequence 012' (S1'-E1', S2'-E2', S3'-E3'), the time points corresponding to S1', S2', S3', and E1', E2', E3' are the time points when the electronic device detects a change in the user's behavior pattern; while in the data segments corresponding to sequence 012 (S1-E1, S2-E2, S3-E3), the time points corresponding to S1, S2, S3, and E1, E2, E3 are the actual time points when the user's behavior pattern changes.

[0177] Combination Figure 12 It can be seen that when electronic devices identify user behavior, the landmark points corresponding to the user's cross-layer movement identified by the electronic devices have a certain error compared to the landmark points of the user's actual real behavior pattern. Specifically, for the S1-E1 segment in sequence 012, the corresponding data segment in the landmark point data identified by the electronic device is the S1'-E2' segment. Compared to the S1-E1 segment, the start time of the S1'-E1' segment is 2 sample points (data frames) ahead, and the end time is 2 sample points (data frames) behind. For the S2-E2 segment in sequence 012, the corresponding data segment in the landmark point data identified by the electronic device is the S2'-E2' segment. Compared to the S2-E2 segment, the start time of the S2'-E2' segment is 2 sample points (data frames) behind, and the end time is 3 sample points (data frames) behind. For the S3-E3 segment in sequence 012, the corresponding data segment in the landmark point data identified by the electronic device is the S3'-E3' segment. Compared to the S3-E3 segment, the start time of the S3'-E3' segment matches, but the end time is 1 sample point (data frame) behind.

[0178] If we use "+" and "-" to represent the predictive model's advance recognition and lag recognition respectively (+ represents advance, i.e., prediction before reality, and - represents lag, i.e. prediction after reality), and use numbers to represent the number of advance or lag sample points (number of data frames), then the error between the landmark point recognition result output by the predictive model (i.e., sequence 012') and the landmark point of the user's actual real behavior pattern (i.e., sequence 012') in the cross-layer motion pattern can be expressed as [+2, -2, 0] (start time) and [-2, -3, -1] (end time).

[0179] Therefore, optionally, during the training process of prediction model 203, the dataset used to train prediction model 203 may also include real landmark data of the user. During the training of prediction model 203, the electronic device can compare the landmark recognition results output by prediction model 203 with the real landmark data of the user to determine the delay time (number of samples) between the start and end times of the predicted landmark and the real landmark, the prediction landmark recognition accuracy (the proportion of landmarks with recognition errors ≤ 3 samples to the total number of landmarks), and the prediction landmark recognition accuracy (the proportion of landmarks with recognition errors ≤ 5 samples to the total number of landmarks). Extensive simulation tests are conducted on erroneous data to continuously optimize the project architecture and model features, ultimately resulting in a more complete behavior pattern recognition algorithm and landmark recognition algorithm.

[0180] It should be noted that, Figure 2 This illustration merely demonstrates the process by which an electronic device identifies a user's behavioral state and does not constitute a limitation on the embodiments of this application. In the actual identification process, the electronic device may perform more or fewer steps. For example, after the data acquisition module 201 completes the data acquisition, the electronic device may perform data completion and data alignment processing on the initial data. This application does not limit this.

[0181] Furthermore, when extracting behavioral semantic features from the aforementioned initial features, the behavioral semantic features extracted by the electronic device (feature engineering module 202) may include one or more of the following: small-amplitude overweight percentage, large-amplitude overweight percentage, small-amplitude weightlessness percentage, large-amplitude weightlessness percentage, continuous small-amplitude overweight percentage, continuous small-amplitude weightlessness percentage, single-sample maximum turning, double-sample maximum turning, mean geomagnetic fluctuation, projection of acceleration onto the direction of gravity, and WIFI change ratio. This application does not limit these features.

[0182] based on Figure 2The provided diagram illustrates the process by which an electronic device identifies a user's behavioral state. This application also provides an architecture diagram of a behavior recognition algorithm, which can be implemented... Figure 2 For details on how the electronic device identifies the user's behavioral state, please refer to [link / reference needed]. Figure 13 .

[0183] like Figure 13 As shown, the algorithm architecture provided in this application includes a data layer, a data processing layer, a core algorithm layer, and an output layer, wherein:

[0184] The data layer, also known as the input layer, is primarily responsible for data acquisition. Specifically, when recognizing user actions, multiple sensors in the electronic device can collect data in real time at a certain frequency, which serves as input to the predictive model; this data can be referred to as initial data. Specifically, the sensors involved in the data layer can include accelerometers, gyroscopes, geomagnetic sensors, and Wi-Fi sensors. Accelerometers, gyroscopes, and geomagnetic sensors can also be referred to as a 3-axis accelerometer, a 3-axis gyroscope, and a 3-axis electronic compass (geomagnetic sensor), respectively; the combination of these three sensors can be called a nine-axis sensor.

[0185] The data processing layer can be responsible for feature extraction from the initial data. Specifically, the extracted features can include statistical features, temporal features, and behavioral semantic features. Optionally, the data processing layer can also be responsible for data alignment and data completion of the initial data.

[0186] The core algorithm layer is primarily responsible for training the prediction model or using the prediction model to perform feature analysis on the data features extracted by the data processing layer, thereby obtaining the behavior recognition results for user actions. Specifically, the prediction model trained by the core algorithm layer and the model that performs feature analysis on the feature data can be XGBoost models.

[0187] The output layer is mainly responsible for outputting the recognition results of the user's actions and behaviors by the above prediction model.

[0188] Specifically, the behavior recognition result can include behavior pattern recognition results. Behavior pattern recognition results can be represented by characters. For example, when the behavior pattern recognition result is a multi-class result (i.e., it needs to specifically identify which behavior pattern the user is in: stationary, walking, going up and down stairs, going up and down elevators, or going up and down escalators), then the recognition result output by the electronic device can be represented by 0, 1, 2, 3, 4, and 5 respectively for the above five behavior patterns. Similarly, when the behavior pattern recognition result is a binary result (i.e., it only needs to identify whether the user is in a horizontal or horizontal movement mode), the recognition result output by the electronic device can be represented by 0 and 1 to indicate the horizontal and horizontal movement modes.

[0189] Furthermore, the aforementioned behavior recognition results may also include landmark point recognition results. Landmark point recognition results can be used to reflect the moment when user behavior switched in the aforementioned behavior pattern recognition results. For example, if a certain sequence of numbers reflects that the user successively performed a horizontal movement mode and a cross-level movement mode, then the landmark point recognition results can reflect at what moment the user ended the horizontal movement and began the cross-level movement.

[0190] Furthermore, the specific functions of the aforementioned data layer, data processing layer, core algorithm layer, and output layer, as well as the content and format of the data involved in each layer, can be found in the aforementioned section. Figures 2-12 The relevant explanations will not be repeated here.

[0191] The recognition process provided in this application extracts behavioral semantic features that better reflect user behavior patterns during the feature extraction stage, and further obtains feature data of different lengths through sliding windows of different scales. This fully combines the differences in details and data of different action patterns to analyze the collected data, which can effectively improve the accuracy of user behavior recognition results. The overall accuracy of the recognition results can reach more than 98%.

[0192] The electronic device provided in the embodiments of this application will be described below.

[0193] The electronic device may be a mobile phone, tablet computer, wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), or dedicated camera (such as SLR camera, point-and-shoot camera), etc. This application does not impose any limitation on the specific type of the electronic device. Specifically, the electronic device may be one of the electronic devices described above.

[0194] Figure 14 The structure of the electronic device is shown as an example.

[0195] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, a Wi-Fi sensor 180N, etc.

[0196] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0197] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0198] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0199] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0200] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0201] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.

[0202] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0203] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0204] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0205] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.

[0206] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0207] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0208] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0209] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0210] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0211] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0212] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0213] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0214] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0215] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0216] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0217] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0218] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0219] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0220] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0221] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0222] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0223] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0224] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0225] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0226] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0227] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0228] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0229] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.

[0230] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.

[0231] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0232] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0233] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0234] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.

[0235] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0236] Geomagnetic sensors, such as the 180D, are mostly used in compasses and map navigation systems. They determine magnetic field strength by measuring changes in resistance. Geomagnetic sensors employ anisotropic magnetoresistance (AMR) materials to detect the magnitude of magnetic field strength in space. This crystalline alloy material is highly sensitive to external magnetic fields; changes in the strength of the magnetic field cause changes in the AMR's resistance. Therefore, when the magnetic field around an electronic device equipped with a geomagnetic sensor changes, even a small change, the sensor can sensitively detect the change in magnetic field strength.

[0237] The accelerometer 180E is a sensor capable of measuring acceleration, typically composed of a mass block, damper, elastic element, sensing element, and adaptation circuitry. When an electronic device is in an accelerating state, the accelerometer 180E can obtain the acceleration value by measuring the inertial force acting on the electronic device and applying Newton's second law. Furthermore, the accelerometer 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the attitude of electronic devices, and is applied in applications such as screen orientation switching and pedometers.

[0238] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.

[0239] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity sensor 180G to detect when a user holds the electronic device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.

[0240] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.

[0241] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0242] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.

[0243] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0244] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.

[0245] A Wi-Fi sensor can be used to search for / scan for existing Wi-Fi network signals and can be used to obtain the physical address of each Wi-Fi signal. It should be understood that the term "Wi-Fi sensor" in this application is used only to refer to a component used for searching for / scanning existing Wi-Fi network signals. The Wi-Fi sensor can be a separate processing element or it can be implemented on the same chip as other elements (such as the other sensors described above). Furthermore, the Wi-Fi sensor can also be stored as program code in the controller's storage element, and its functions can be called and executed by a processing element of the processor.

[0246] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0247] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0248] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0249] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0250] In some embodiments, the prediction model provided in the present application embodiments may be deployed in the internal memory 121 of the electronic device 100 or the memory in the processor 110, such as the prediction model 203 described above.

[0251] In some embodiments, the accelerometer 180E, gyroscope 180B, geomagnetic sensor 180D, and WIFI sensor 180D in the electronic device 100 can collect data in real time. The electronic device can then extract features from the data collected by these sensors to obtain temporal feature data, statistical feature data, and behavioral semantic feature data (the specific content and meaning of the data can be found in Table 1 above). Subsequently, the electronic device can input the aforementioned temporal feature data, statistical feature data, and behavioral semantic feature data into the aforementioned prediction model. Since the prediction model performs a series of complex calculations such as feature extraction on these feature data, it finally outputs the behavior recognition result of the user's behavior.

[0252] In some embodiments, after acquiring the initial data collected by the sensor, the electronic device can segment the initial data using a sliding window with a sampling window length of 2.56s and a window coverage of 50%, to obtain the first data. Further, in some embodiments, during feature extraction, the electronic device can simultaneously segment the first data using sample windows of different sizes, obtaining multiple sets of data with varying data segment lengths (these multiple sets of data can be referred to as target data). During feature extraction, features are extracted from each of these multiple sets of data, thereby effectively improving the performance of the prediction model and the accuracy of the recognition results. For example, when determining the size of the sample window for the first data, the specific size of the sample window can be set to 5, 10, and 15, respectively.

[0253] In some embodiments, the behavior recognition results output by the above prediction model may include behavior pattern recognition results. Behavior pattern recognition results can be represented by characters. For example, when the behavior pattern recognition result is a multi-class result (i.e., it is necessary to specifically identify which behavior mode the user is in: stationary, walking, climbing stairs, using an elevator, or using an escalator), the recognition results output by the electronic device can be represented by 0, 1, 2, 3, 4, and 5, respectively. Similarly, when the behavior pattern recognition result is a binary result (i.e., it is only necessary to identify whether the user is in a level-based movement mode or a cross-level movement mode), the recognition results output by the electronic device can be represented by 0 and 1, indicating level-based and cross-level movement modes, respectively. Furthermore, the above behavior recognition results may also include landmark point recognition results. Landmark point recognition results can be used to reflect the moment when the user's behavior switches in the above behavior pattern recognition results. For example, if a certain sequence of numbers reflects that the user successively performed level-based and cross-level movement modes, the landmark point recognition results can reflect at what moment the user ended level-based movement and began cross-level movement.

[0254] Figure 15This is a flowchart illustrating an identification method provided in an embodiment of this application. The method extracts behavioral semantic features from data collected by sensors in an electronic device that better reflect user behavior patterns, and combines these behavioral semantic features to identify user behavior, effectively improving the accuracy of user behavior identification results. Figure 15 As shown, the method provided in this application embodiment may include the following steps:

[0255] S101. The electronic device extracts features from the target data to obtain feature data, which includes behavioral semantic feature data.

[0256] The aforementioned electronic device may be the electronic device 100 described above.

[0257] The aforementioned target data can be obtained based on initial data collected by sensors in the aforementioned electronic device. The specific process involved can be found in the descriptions of the foregoing and subsequent embodiments, and will not be repeated here. In this method, the electronic device may include multiple sensors, such as accelerometers, gyroscopes, geomagnetic sensors, and Wi-Fi sensors. These sensors can collect corresponding sensor data at a certain frequency, and this data constitutes the aforementioned initial data.

[0258] Feature extraction refers to the process of transforming the aforementioned target data into feature data that can characterize user behavior. In this application, in order to focus on studying human behavior patterns and accurately identify five behaviors—stationary, walking, climbing stairs, escalator, and elevator—the electronic device, during the feature extraction process of the aforementioned target data, combines the differences in details and data exhibited by various user behavior patterns to extract the aforementioned behavioral semantic feature data from the target data. This behavioral semantic feature data includes first feature data characterizing the electronic device in a state of hypergravity / weightlessness, and / or second feature data characterizing changes in the orientation of the electronic device (user).

[0259] Furthermore, the behavioral semantic features extracted from electronic devices may also include one or more features reflecting changes in the geomagnetic intensity of the environment in which the electronic device is located, changes in the magnitude of the acceleration of the electronic device (user) in the direction of gravity, and changes in the number of WIFI networks in the environment in which the electronic device (user) is located.

[0260] In addition to the aforementioned behavioral semantic features, electronic devices can also extract physically meaningful temporal features and statistically significant features based on the target data. In other words, the feature data extracted by electronic devices from the target data can be broadly categorized into three types: statistical feature data, temporal feature data, and behavioral semantic features. The specific content and meaning of each of these three categories of feature data can be found in the preceding explanation of Table 1, and will not be repeated here.

[0261] It should be noted that, in an optional implementation, among the first, second, third, fourth, and fifth feature data included in the above-mentioned behavioral semantic features, the electronic device may simultaneously extract all data features from these data features, or it may extract only some of the data features from these data features; this application does not limit this. For example, the electronic device may not extract the first and second feature data, but only extract the third feature data. Compared to extracting only the statistical feature data and temporal feature data, the behavior recognition result obtained by the electronic device combining the third feature data can also achieve a certain degree of accuracy.

[0262] Similarly, in an optional implementation, among the 11-dimensional behavioral semantic features involved in the first, second, third, fourth, and fifth feature data, such as the proportion of slight overweight, significant overweight, slight weightlessness, significant weightlessness, continuous slight overweight, continuous slight weightlessness, single-sample maximum turning, double-sample maximum turning, mean geomagnetic fluctuation, projection of acceleration onto the direction of gravity, and WIFI change ratio, the electronic device may extract only one or more dimensions of the behavioral semantic features, which can also improve the accuracy of the behavior recognition results to a certain extent.

[0263] S102. The electronic device performs feature analysis on the above feature data to obtain behavior recognition results.

[0264] The aforementioned behavior recognition results include a first recognition result, which represents the user's current behavior pattern.

[0265] Specifically, the aforementioned electronic device can be equipped with a predictive model that performs feature analysis on the aforementioned feature data. This predictive model can be the predictive model 203 described above, specifically, it can be an XGBoost model. The training dataset used for training this predictive model can be sensor data collected by multiple devices through their respective sensors (e.g., the accelerometer, gyroscope, magnetometer, and Wi-Fi sensor described above). The manufacturer of the aforementioned electronic device can use the data set collected by the aforementioned multiple devices as the training set for the predictive model, and then use a device with model training capabilities to train the predictive model. Subsequently, during the production process of the aforementioned electronic device, the predictive model is uniformly deployed in the electronic device. Then, when the aforementioned electronic device with the predictive model deployed identifies user behavior (i.e., during the use of the predictive model), the aforementioned electronic device can use only the sensor data collected by its own sensors (i.e., the aforementioned initial data) as input to the predictive model, or use the data obtained by preprocessing the initial data (e.g., the aforementioned target data) as input to the predictive model to obtain the recognition result of user behavior.

[0266] Specifically, the aforementioned behavioral patterns can be categorized into level-based behavioral patterns and cross-level behavioral patterns. Level-based behavioral patterns can include stationary behavior and walking, while cross-level behavioral patterns can include climbing stairs, using elevators, and using escalators. Optionally, when predicting user behavioral patterns, the prediction model can employ a multi-class to binary classification approach. It first determines which of the five behavioral patterns (stationary, walking, climbing stairs, using elevators, and using escalators) the user's action belongs to, then processes the results according to requirements, merging the identification results. Specifically, stationary behavior and walking are categorized as level-based behavioral patterns, while climbing stairs, using elevators, and using escalators are categorized as cross-level behavioral patterns.

[0267] Furthermore, the behavior recognition results output by the prediction model 203 can represent the user's behavior pattern using characters. Specifically, it can be represented as a sequence of numbers, where each number represents a duration during which the user acted in the corresponding behavior pattern. This duration can be the duration of a data frame. For example, when the behavior pattern recognition result is a multi-class result (i.e., it is necessary to specifically identify which behavior pattern the user is in: stationary, walking, climbing stairs, using an elevator, or using an escalator), the recognition results output by the electronic device can be represented by 0, 1, 2, 3, 4, and 5, respectively. Similarly, when the behavior pattern recognition result is a binary result (i.e., it is only necessary to identify whether the user is in a horizontal or horizontal movement pattern), the recognition results output by the electronic device can be represented by 0 and 1, respectively, to indicate the horizontal and horizontal movement patterns.

[0268] For example, suppose that after the prediction model performs feature analysis on the feature data, the final output behavior pattern recognition result is a sequence of numbers such as [0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 0, 0]. This sequence of numbers represents the user moving in a flat motion mode for a period of time, which is equal to the duration of 5 data frames. Then, the user moves in a cross-layer motion mode for another period of time, which is equal to the duration of 6 data frames. Finally, the user moves in a flat motion mode for another period of time, which is equal to the duration of 2 data frames.

[0269] Optionally, in some embodiments, the behavior recognition result may further include the user's landmark recognition result. This landmark recognition result characterizes the point in time when the user switched behavior modes. In some embodiments, this landmark recognition result may be referred to as a second recognition result. That is, the landmark recognition result can be used to reflect the moment when the user's behavior switched in the aforementioned behavior mode recognition result. For example, if a certain sequence of numbers reflects that the user successively performed a horizontal movement mode and a cross-level movement mode, then the landmark recognition result can reflect at what moment the user ended the horizontal movement and began the cross-level movement. For details, please refer to the foregoing... Figure 12 The relevant explanations will not be repeated here.

[0270] It should be understood that the landmark recognition results described above are predictions of user behavior made by electronic devices based on data collected by sensors. However, there may still be some time discrepancies between these predictions and the actual landmarks. Therefore, optionally, the dataset used to train the prediction model can also include real landmark data of the user. During the training of the prediction model, the electronic device can compare the landmark recognition results output by the prediction model with the user's real landmark data to determine the delay time (number of samples) between the start and end times of the predicted landmarks and the real landmarks, the prediction landmark recognition accuracy (the percentage of landmarks with recognition errors ≤ 3 samples out of the total number of landmarks), and the prediction landmark recognition accuracy (the percentage of landmarks with recognition errors ≤ 5 samples out of the total number of landmarks). Extensive simulation tests are conducted on the erroneous data to continuously optimize the project architecture and model features, ultimately resulting in a more complete behavior pattern recognition algorithm and landmark recognition algorithm.

[0271] As explained above, in the identification method provided in the application, the initial data collected by the sensors in the electronic device needs to be processed multiple times before it can be used as feature data for training. The following will combine... Figure 16 The specific process by which the electronic device obtains the aforementioned feature data based on the initial data is explained.

[0272] Figure 16 A flowchart illustrating a data processing method provided in this application. Figure 16 As shown, the method may include:

[0273] S201. Electronic equipment collects initial data.

[0274] The aforementioned electronic device may be the electronic device 100 described above.

[0275] In this method, the electronic device may include multiple sensors, such as an accelerometer, a gyroscope, a geomagnetic sensor, and a Wi-Fi sensor. These sensors can collect corresponding sensor data at a certain frequency, and this data is the initial data mentioned above.

[0276] Specifically, each sensor in the electronic device can collect data at a certain sampling frequency, such as 100Hz, to obtain the aforementioned initial data.

[0277] S202. The electronic device uses a sliding window of the first length to segment the initial data to obtain the first data.

[0278] It should be understood that since actions take a certain amount of time to occur, the data collected by the sensor at a single moment is insufficient to reflect the characteristics of the user's actions. Therefore, after obtaining the initial data, the electronic device can select a sliding window of a certain length and coverage to segment the initial data collected by the sensor. That is, the initial data of the long time series is divided into smaller data frames (the length of each data frame is the length of the sliding window) for processing, so as to facilitate feature extraction.

[0279] It is important to understand that the length of the data frame, i.e., the length of the sliding window, directly affects not only the quality of feature extraction and classification performance, but also whether the method is suitable for real-time systems. For example, if the sliding window is too small, each segmented data frame may only represent a part of the action, failing to reflect the complete behavioral state of the action. This results in the extracted features not effectively representing the action, leading to a decrease in recognition performance, and the electronic device needs to perform frequent recognition, increasing the computational load. Conversely, if the sliding window is too large, a single data frame may contain multiple actions, causing the system to be unable to effectively recognize the current behavior, affecting the system's real-time performance, and also impacting the accuracy of the recognition results. Therefore, when segmenting the initial data, the length of the sliding window directly affects the accuracy of the recognition results.

[0280] Current research indicates that a window length of 2.56 s is ideal when the sampling frequency is 100 Hz. This is because Fourier transform is required for frequency domain calculations, and using a power of 2 as the time window ensures the integrity of the data involved in the calculations. Therefore, the aforementioned first length can optionally be 2.56 s. Of course, the aforementioned first length can also be other values, and this application does not limit it.

[0281] In addition, the sliding window coverage is also an important factor affecting the accuracy of the recognition results. The sliding window coverage refers to the overlap rate between two adjacent segmentation windows during the data segmentation process. A coverage rate of 0% means that adjacent windows do not overlap, and a coverage rate of 50% means that the current window contains half of the data from the previous window.

[0282] In this embodiment, the electronic device can set the window coverage of the sliding window to 50%. Extensive experimental data shows that setting the window coverage to 50% can effectively reduce the disturbance caused by transitional behavior. Therefore, in this embodiment, the electronic device can use a sliding window with a window length of 2.56s and a window coverage of 50% to segment the initial data. For each sensor, the data it collects will be segmented into multiple data segments, one of which can be called a "data frame". It is easy to see that when the electronic device samples at a frequency of 100Hz and uses a sliding window with a window length of 2.56s and a window coverage of 50% to segment the initial data collected by the sensor, each data frame (which can also be called a sample during model training and testing) contains data from 256 time points. For details, please refer to the aforementioned... Figure 3 The relevant descriptions will not be repeated here.

[0283] Optionally, before performing data segmentation on the initial data, the electronic device may perform a series of preprocessing operations on the initial data, such as data alignment, data completion, noise reduction, filtering, etc. After completing these preprocessing operations, the obtained data is segmented using the sliding window of the first length.

[0284] S203. The electronic device uses at least two sample windows to segment the first data to obtain at least two sets of data fragments, and uses the at least two sets of data fragments as the target data.

[0285] As explained above, samples extracted from a single data frame cannot adequately reflect the semantic features before and after the current frame. For example, during an elevator's operation, there are periods of constant speed. If only a single data frame is used to extract features at this moment, the features will not differ significantly from those at a stationary position. In other words, if only a single data frame is used for feature extraction, the information captured within the window is too limited, potentially leading to low accuracy in the prediction model's output. Therefore, the aforementioned electronic device can segment the first data using sample windows based on data frames, and extract features in batches of several data frames during the feature extraction process.

[0286] Furthermore, the duration of the relevant feature data varies for different behavior patterns. For example, when a user is moving between floors in an elevator, the sample window needs to be longer to fully extract feature data indicating whether the user is in a state of weightlessness or weightlessness. However, when the user is in a stationary or moving mode on a level floor, only a few data frames are needed to reflect the user's motion characteristics, so the sample window can be set shorter accordingly. Therefore, in some embodiments, the electronic device can simultaneously set multiple sample windows of different sizes to segment the first data.

[0287] Therefore, in this method, after obtaining the first data, the electronic device can use at least two sample windows to segment the first data to obtain at least two sets of data fragments, and use the at least two sets of data fragments as the target data.

[0288] It is important to note that, unlike the segmentation of initial data collected by various sensors using a sliding window, the size of the sample window used for segmenting the first data is determined by the data frame as the smallest unit, and the electronic device can simultaneously use multiple sample windows of different sizes to segment the first data. However, when segmenting the initial data, the window size is determined by the number of time points as the smallest unit, and the coverage between windows can be set to 50%. For example, if the electronic device samples the initial data at a frequency of 100Hz, and the window size determined by the electronic device for the initial data is 2.56s, then when segmenting the initial data, the number of data points in each window is 256. These 256 data points correspond to 256 sampling time points, and each data segment obtained from segmenting the initial data is a data frame. All data frames obtained from the initial data constitute the first data. In subsequent segmentations of the first data, the electronic device can determine the sample window for the first data in units of data frames.

[0289] Specifically, the aforementioned at least two sample windows can be three sample windows, and the aforementioned at least two sets of data segments constitute three sets of data segments. Furthermore, the specific sizes of the three sample windows can be set to 5 data frame lengths, 10 data frame lengths, and 15 data frame lengths, respectively. These window sizes were determined through observation of a large amount of experimental data. In the aforementioned electronic device, with each sensor sampling frequency of 100Hz, a data frame length of 2.56s (one data frame length), and three sample windows, the prediction model achieves the best accuracy when the sample window sizes are set to 5, 10, and 15 data frame lengths, with an output accuracy of approximately 0.98. For details, please refer to the aforementioned... Figure 11 The relevant explanations will not be repeated here.

[0290] S204. The electronic device extracts features from the target data to obtain feature data.

[0291] After obtaining the target data, the electronic device can perform feature extraction on the target data to obtain the feature data.

[0292] The feature data extracted by electronic devices from the aforementioned target data can be broadly categorized into three types: statistical feature data, temporal feature data, and behavioral semantic features. The specific content and meaning of each of these three categories of feature data can be found in the preceding explanation of Table 1, and will not be repeated here.

[0293] This application also provides an electronic device, which includes one or more processors and a memory; wherein the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the methods shown in the foregoing embodiments.

[0294] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0295] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0296] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A recognition method, characterized in that, The method includes: The initial data is segmented using a sliding window of the first length to obtain the first data, which includes multiple data frames of the first length. The first data is segmented using at least two sample windows to obtain at least two sets of data fragments, wherein the length of any sample window in the at least two sample windows is different from the length of the other windows in the at least two sample windows. Using the at least two sets of data fragments as target data, feature extraction is performed on the target data to obtain feature data. The feature data includes behavioral semantic feature data. The target data is obtained based on initial data collected by sensors in the electronic device. The behavioral semantic feature data includes first feature data characterizing the electronic device in a state of hypergravity / weightlessness, and / or second feature data characterizing the change in orientation of the electronic device. Feature analysis is performed on the feature data to obtain behavior recognition results. The behavior recognition results include a first recognition result and a second recognition result. The first recognition result represents the user's current behavior pattern, and the second recognition result represents the landmark point of the moment when the user switches behavior patterns.

2. The method according to claim 1, characterized in that, The step of segmenting the first data using at least two sample windows to obtain at least two sets of data fragments includes: The first data is segmented using three sample windows to obtain three sets of data segments. The window lengths of the three sample windows are 5 times the first length, 10 times the first length, and 15 times the first length, respectively.

3. The method according to claim 1 or 2, characterized in that, The behavioral semantic feature data further includes at least one of the third feature data, the fourth feature data, and the fifth feature data, wherein: The third feature data characterizes the changes in geomagnetic intensity in the environment detected by the electronic device; The fourth feature data characterizes the change in the magnitude of the acceleration of the electronic device in the direction of gravity; The fifth feature data characterizes the change in the number of WIFI networks detected by the electronic device.

4. The method according to claim 1 or 2, characterized in that, The first feature data includes at least one of the following: slight weightlessness / weight gain percentage, significant weightlessness / weight gain percentage, and continuous slight weightlessness / weight gain percentage. The slight weightlessness / weight gain percentage represents the percentage of times in the first data segment when the electronic device is in a weightlessness / weight gain state and the weightlessness / weight gain value is less than a second threshold, within the total number of times in the first data segment. The significant weightlessness / weight gain percentage represents the percentage of times in the first data segment when the electronic device is in a weightlessness / weight gain state and the weightlessness / weight gain value is greater than a third threshold, within the total number of times in the first data segment. The percentage of continuous small-amplitude weightlessness / weightlessness represents the proportion of the number of moments corresponding to the maximum duration of continuous weightlessness / weightlessness of the electronic device in the first data segment to the total number of moments in the first data segment; the first data segment is any data segment in the target data.

5. The method according to claim 3, characterized in that, The first feature data includes at least one of the following: slight weightlessness / weight gain percentage, significant weightlessness / weight gain percentage, and continuous slight weightlessness / weight gain percentage. The slight weightlessness / weight gain percentage represents the percentage of times in the first data segment when the electronic device is in a weightlessness / weight gain state and the weightlessness / weight gain value is less than a second threshold, within the total number of times in the first data segment. The significant weightlessness / weight gain percentage represents the percentage of times in the first data segment when the electronic device is in a weightlessness / weight gain state and the weightlessness / weight gain value is greater than a third threshold, within the total number of times in the first data segment. The percentage of continuous small-amplitude weightlessness / weightlessness represents the proportion of the number of moments corresponding to the maximum duration of continuous weightlessness / weightlessness of the electronic device in the first data segment to the total number of moments in the first data segment; the first data segment is any data segment in the target data.

6. The method according to claim 1 or 2, characterized in that, The second feature data includes at least one of single-sample maximum turning or double-sample maximum turning, wherein the single-sample maximum turning represents the user's maximum turning angle during the sampling duration corresponding to a data frame; The dual-sample maximum steering characterizes the maximum sum of the steering angles of a user at any two consecutive sampling moments in a data frame.

7. The method according to claim 3, characterized in that, The second feature data includes at least one of single-sample maximum turning or double-sample maximum turning, wherein the single-sample maximum turning represents the user's maximum turning angle during the sampling duration corresponding to a data frame; The dual-sample maximum steering characterizes the maximum sum of the steering angles of a user at any two consecutive sampling moments in a data frame.

8. The method according to claim 3, characterized in that, The third feature data includes the mean geomagnetic fluctuation, which represents the average value of the degree of data change corresponding to multiple geomagnetic data segments of the second length. The degree of data change corresponding to any geomagnetic data segment of the multiple geomagnetic data segments of the second length is determined by the maximum and minimum geomagnetic values ​​in any geomagnetic data segment. The multiple geomagnetic data segments of the second length are all included in the first data segment, which is any data segment in the target data.

9. The method according to claim 3, characterized in that, The fifth feature data includes the WIFI change ratio, which represents the ratio of the first WIFI network number to the second WIFI network number. The second WIFI network number is the number of WIFI network addresses contained in the last data frame of the first data segment, and the first WIFI network number is the number of WIFI network addresses that exist in the last data frame of the first data segment but do not exist in the first data frame of the first data segment.

10. An electronic device, characterized in that, The electronic device includes: one or more processors, memory, and a display screen; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-9.

11. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-9.

12. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Special population-oriented danger sensing and alarming system

    CN103530978A

  • Hybrid floor positioning method based on floor switching behavior recognition

    CN109579846A

  • Human body action recognition method based on dual-channel residual neural network

    CN110348494A