Wearable micro-grid energy optimization distribution method based on dynamic environment perception

By integrating flexible piezoelectric fibers and photovoltaic units into wearable devices, combining dynamic environmental perception and deep reinforcement learning, and optimizing energy distribution, the inefficiency of existing wearable devices in energy management and thermal management is solved, achieving efficient energy utilization and precise temperature control.

CN120675188AInactive Publication Date: 2025-09-19KUNSHAN YUNJING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510791749.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing wearable devices suffer from inefficiency and lack of precision in energy and thermal management, resulting in energy waste and a poor user experience.

Method used

A wearable microgrid energy optimization distribution method based on dynamic environmental perception is adopted. User motion and ambient light data are collected through flexible piezoelectric fibers and flexible photovoltaic units. A dual-channel neural network is combined to predict user behavior and environmental changes. A deep reinforcement learning algorithm is used to optimize energy distribution and achieve millisecond-level temperature regulation.

Benefits of technology

It improves energy utilization efficiency, extends equipment usage time by 30-50%, enhances user comfort, achieves temperature control accuracy of ±0.5℃, and improves the system's adaptability in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675188A_ABST
    Figure CN120675188A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wearable equipment, in particular to a wearable micro-grid energy optimal distribution method based on dynamic environment perception, and the method comprises the steps: collecting user motion data and environment illumination data through flexible piezoelectric fibers arranged at human joints and a flexible photovoltaic unit arranged on the outer surface of clothes; a two-channel neural network is utilized to predict a motion mode and illumination intensity change of a user at a next moment, a human body thermodynamics constraint model is established based on human body thermal characteristics, and a heat supply mode and heat supply electric quantity are determined through a deep reinforcement learning algorithm. And then energy distribution control is carried out on the wearable micro-grid, the supercapacitor and the flexible fiber battery through a gradient game priority scheduling mechanism, millisecond-level temperature adjustment is realized, the method combines environment perception and a deep reinforcement learning algorithm, an energy distribution strategy can be dynamically optimized, and under the same energy input condition, the service time of equipment is prolonged by 30-50%, and the service life of the equipment is prolonged by 30-50%. And the energy utilization efficiency is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wearable devices, and in particular to a wearable microgrid energy optimization distribution method based on dynamic environment perception, which is used for intelligent management of energy distribution and temperature control of wearable devices. Background Art

[0002] Wearable devices are increasingly used in daily life, healthcare, and professional fields. However, existing wearable devices face challenges with inefficient energy management and imprecise thermal management. Traditional wearable devices typically rely on a single energy source, such as batteries, which makes them difficult to maintain for extended periods of use. Furthermore, existing thermal management technologies often employ fixed control strategies or simple feedback control, failing to optimize based on user activity and environmental changes. This results in energy waste and a poor user experience.

[0003] Existing technologies typically employ static preset schemes or simple feedback control strategies for energy management, which cannot effectively respond to complex and changing usage environments and user needs. Furthermore, traditional thermal management devices often use rigid materials, resulting in poor wearer comfort, uneven heat distribution, and inability to precisely control temperature. Furthermore, energy harvesting and energy management technologies are often disconnected, lacking overall coordinated optimization and failing to fully utilize multiple energy sources and intelligent control algorithms to improve energy efficiency.

[0004] Therefore, there is an urgent need for a wearable microgrid method that can optimize energy distribution based on dynamic environmental perception to improve energy utilization efficiency and user experience. Summary of the Invention

[0005] The purpose of this invention is to provide a wearable microgrid energy optimization allocation method based on dynamic environment perception, aiming to solve the problems of low energy management efficiency, inaccurate thermal management and poor user experience in the existing technology.

[0006] The present invention proposes a wearable microgrid energy optimization allocation method based on dynamic environment perception, including:

[0007] Obtain user motion data and ambient lighting data, including:

[0008] The user's motion data is collected through flexible piezoelectric fibers set at the human joints;

[0009] Ambient light data is collected through flexible photovoltaic units installed on the outer surface of clothing;

[0010] Establish a dynamic environment prediction model and a human body thermodynamic constraint model, including:

[0011] Based on the user motion data and the ambient light data, a dual-channel neural network is used to predict the user's motion pattern and light intensity change at the next moment, respectively, to obtain user motion prediction data and light prediction data;

[0012] Establish a human body thermodynamic constraint model based on the thermal characteristics of the human body, including the human body thermal conductivity model, human body heat dissipation model, human body metabolic heat production model and thermal balance model;

[0013] Determine energy optimization allocation strategies, including:

[0014] Determining a heating mode and heating power using a deep reinforcement learning algorithm based on the user motion prediction data, the light prediction data, and the human body thermodynamic constraint model;

[0015] According to the heating mode and the heating power, energy distribution control is performed on the wearable microgrid, supercapacitor and flexible fiber battery through a gradient game priority scheduling mechanism to achieve millisecond-level temperature regulation.

[0016] Preferably, the dual-channel neural network comprises:

[0017] Motion prediction channel, used to receive user motion pattern and time step data and output the user's future motion pattern sequence;

[0018] Light prediction channel, which receives current light intensity and time data and outputs future light intensity sequence;

[0019] Among them, the motion prediction channel and the illumination prediction channel both adopt a long short-term memory neural network structure, and introduce an attention weight mechanism to improve prediction accuracy.

[0020] Preferably, the establishment of the human body thermodynamic constraint model includes:

[0021] Human thermal conductivity model: describes the relationship between human metabolic heat production and skin thermal conductivity, skin-air contact area, and temperature gradient;

[0022] Human body heat dissipation model: describes the relationship between the heat exchange between the human body and the air, the skin surface heat release coefficient, the ambient temperature and the skin temperature;

[0023] Human metabolic heat production model: describes the relationship between human metabolic heat production, metabolic power and total heat production rate;

[0024] Thermal balance model: describes the relationship between the rate of change of skin temperature and skin mass, heat ratio, metabolic heat, radiation and heat exchange between the human body and the air.

[0025] Preferably, the heating mode includes:

[0026] No heating mode: Select this mode when the current body temperature has reached the standard or the user's activities generate enough heat;

[0027] Skin heating mode: Select when the ambient temperature is low but not extreme and slight heating is required;

[0028] Muscle heating mode: Select this mode when the ambient temperature is low and the user's activity level is low, and deep heating is required;

[0029] Full heating mode: Select this mode when the environment is extremely cold or the user has special needs and needs to heat both the skin and muscles;

[0030] Among them, the selection of heating mode is determined based on the real-time body temperature model, heating mode classification model and heating amount model.

[0031] Preferably, the state space of the deep reinforcement learning algorithm includes: the user's real-time motion pattern; the predicted light intensity; the wearable microgrid input power; the real-time skin temperature; the action space of the deep reinforcement learning algorithm includes: the maximum power supply time under the user's real-time motion state; the maximum power supply time of the energy collection layer;

[0032] The system is guided to learn the optimal strategy through a profit function, which evaluates the difference between the predicted value and the measured value and the temperature difference between the real-time body temperature and the set temperature value.

[0033] Preferably, the deep reinforcement learning algorithm adopts an improved network structure to decompose the action value function into a state value part and an action advantage part, wherein:

[0034] The state value component assesses the intrinsic value of the current state;

[0035] The action advantage component assesses the relative advantage of taking a specific action in a specific state;

[0036] Improve learning efficiency and policy stability by estimating state value and action advantage separately.

[0037] Preferably, the gradient game priority scheduling mechanism includes:

[0038] Evaluate the priority of each energy unit based on energy availability, efficiency factor, user preference and timeliness;

[0039] Treat each energy unit as a game participant and reach the equilibrium point of resource allocation through iterative calculation;

[0040] Evaluate the current strategy based on actual results and adjust the allocation parameters along the gradient direction;

[0041] After multiple iterations, the optimal energy allocation strategy is converged to achieve dynamic coordination among the components of the energy system.

[0042] Preferably, the millisecond-level temperature adjustment specifically includes:

[0043] Execute 15 power switches in 0.5 seconds to achieve fine control of the heating module;

[0044] Determine the optimal heating mode based on the current heat demand and body temperature difference;

[0045] Calculate required power levels and determine power switching parameters;

[0046] The charging and discharging states of the energy storage unit including the supercapacitor and the flexible fiber battery are controlled according to the power switching parameter. When the power parameter is positive, the supercapacitor and the flexible fiber battery are discharged, and when the power parameter is negative, the supercapacitor and the flexible fiber battery are charged.

[0047] Preferably, the energy management of the flexible piezoelectric fiber and the flexible photovoltaic unit includes:

[0048] When the ambient light is sufficient, the energy collected by the flexible photovoltaic unit is used first;

[0049] When the user's activity intensity is high, the energy harvested by the flexible piezoelectric fiber is preferentially used;

[0050] When the energy demand is greater than the current harvested energy, supplementary energy is provided by supercapacitor arrays and flexible fiber batteries;

[0051] When the harvested energy exceeds the energy demand, the remaining energy is stored in the supercapacitor array and flexible fiber battery;

[0052] Among them, energy allocation decisions are dynamically adjusted based on predicted future energy demand and energy collection conditions.

[0053] Preferably, the flexible piezoelectric fiber and the flexible photovoltaic unit are implemented as follows:

[0054] The flexible piezoelectric fiber is made of flexible nano-piezoelectric material and is connected to the wearable microgrid via flexible electrodes;

[0055] The flexible photovoltaic unit is prepared by plasma enhanced chemical vapor method using amorphous silicon material on a polyimide substrate, and finally flexibly encapsulated with ethylene-tetrafluoroethylene copolymer;

[0056] The supercapacitor uses an interdigitated structure made of MXene ink and an electrolyte made of polyvinyl alcohol-based hydrogel, and the 3D interdigitated structure is formed by 3D printing technology;

[0057] The flexible fiber battery combines MXene materials with LiMn2O4 (positive electrode) and Li4Ti5O 12(Anode) composite, prepared by microfluidic wet spinning technology;

[0058] Among them, the application of the flexible material improves the wearing comfort and energy collection efficiency of the device.

[0059] The present invention has the following beneficial effects:

[0060] 1. Improve energy efficiency: By combining environmental perception and deep reinforcement learning algorithms, the present invention can dynamically optimize energy allocation strategies and extend the device usage time by 30-50% under the same energy input conditions.

[0061] 2. Enhanced user comfort: Based on a precise human body thermodynamic model and millisecond-level temperature control, this invention can improve temperature control accuracy to ±0.5°C, significantly better than the ±2°C accuracy of existing products, providing a more comfortable wearing experience.

[0062] 3. Improved system adaptability: Through a dual-channel prediction mechanism and deep reinforcement learning, the system can automatically adapt to different environments and user states, and maintain stable performance in various environments with a temperature range of -10°C to 40°C.

[0063] 4. Optimize the wearing experience: Using flexible materials and advanced manufacturing processes to make the device lighter, thinner and more comfortable while maintaining high performance.

[0064] 5. Achieve intelligent control: Through the collaboration of edge computing and real-time control, the system response delay is reduced to less than 100ms, achieving true real-time intelligent control. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 Schematic diagram of the system architecture of the wearable microgrid energy optimization distribution method based on dynamic environment perception of the present invention;

[0066] Figure 2 It is a schematic diagram of the dual-channel neural network structure of the present invention;

[0067] Figure 3 Schematic diagram of the composition of the human body thermodynamic constraint model in the present invention;

[0068] Figure 4 It is a schematic diagram of the relationship between the state space and action space of deep reinforcement learning in the present invention;

[0069] Figure 5 It is a workflow diagram of the gradient game priority scheduling mechanism in the present invention;

[0070] Figure 6 It is a timing diagram of millisecond-level temperature regulation in the present invention;

[0071] Figure 7It is a schematic structural diagram of the flexible piezoelectric fiber and the flexible photovoltaic unit in the present invention. DETAILED DESCRIPTION

[0072] Please refer to the attached Figure 1-7 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood by those skilled in the art that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0073] like Figure 1 As shown, the present invention provides a wearable microgrid energy optimization distribution method based on dynamic environment perception, comprising the following steps:

[0074] 1. Obtain user motion data and ambient lighting data.

[0075] In a preferred embodiment of the present invention, the system collects user motion data through flexible piezoelectric fibers installed at human joints. The flexible piezoelectric fibers are made of flexible nano-piezoelectric materials and are typically 50-100 microns thick. This thickness ensures sufficient flexibility while maintaining good piezoelectric properties. The piezoelectric fibers are installed at joints such as the elbows and knees. When the joints bend, the fibers deform and generate a voltage signal. Preferably, 3-5 pieces of piezoelectric fiber cloth are installed at each joint to ensure that motion information is captured from different directions.

[0076] At the same time, the system collects ambient light data through flexible photovoltaic cells installed on the outer surface of the clothing. These photovoltaic cells are made of low-emissivity glass and silicon semiconductor materials through a nanoimprint process. Their thickness is typically controlled between 80 and 150 microns, and their area can be adjusted between 400 and 800 square centimeters depending on the clothing design. The photovoltaic cells not only convert light energy into electrical energy, but also reflect the ambient light intensity through changes in output voltage. Generally, for every 100 lux increase in light intensity, the output voltage increases by approximately 0.2 to 0.5 volts.

[0077] In practical applications, flexible piezoelectric fibers and flexible photovoltaic cells simultaneously collect data, typically at a sampling rate of 1Hz in low-activity states and 10Hz in high-activity states to balance data accuracy and energy consumption. The collected signals are transmitted via flexible conductive lines to edge computing nodes for preliminary digital signal processing.

[0078] 2. Establish a dynamic environment prediction model and a human body thermodynamic constraint model.

[0079] Based on the user's motion data and ambient lighting data, the system further establishes a dynamic environment prediction model and a human body thermodynamic constraint model.

[0080] like Figure 2As shown, the system utilizes a dual-channel neural network to predict the user's next motion pattern and light intensity changes, respectively. This dual-channel network comprises a motion prediction channel and a light prediction channel. The motion prediction channel receives the user's motion pattern and time step data and outputs a sequence of the user's motion patterns for the next 0.3 to 300 seconds. The light prediction channel receives current light intensity and time data and outputs a sequence of future light intensities. Preferably, both channels utilize a long-short-term memory neural network architecture. To improve prediction accuracy, the system incorporates an attention weighting mechanism, assigning different weights to predictions at different time steps. For example, the weight of a near-term prediction (t+0.1 seconds) can be set to 0.6, the weight of a medium-term prediction (t+0.2 seconds) to 0.3, and the weight of a long-term prediction (t+0.3 seconds) to 0.1. This weighting helps the system focus on more important time points, improving overall prediction accuracy.

[0081] At the same time, if Figure 3 As shown, the system establishes a human body thermodynamic constraint model based on the thermal characteristics of the human body. This model consists of four sub-models: a human thermal conductivity model, a human heat dissipation model, a human metabolic heat production model, and a thermal balance model. The human thermal conductivity model describes the relationship between human metabolic heat production and the skin's thermal conductivity, the skin-air contact area, and the temperature gradient. In practical applications, the skin's thermal conductivity is typically between 0.2 and 0.5 W / m·Kelvin, and the contact area is adjusted between 1.5 and 2.2 square meters based on the user's body shape. The human heat dissipation model describes the relationship between the amount of heat exchange between the human body and the air, the skin's surface heat dissipation coefficient, the ambient temperature, and the skin temperature. The skin's surface heat dissipation coefficient is typically between 4 and 6 W / m·Kelvin. The human metabolic heat production model describes the relationship between metabolic power and total heat production rate. Metabolic power is approximately 60-80 watts at rest and can reach 400-600 watts during intense exercise. The thermal balance model comprehensively considers these factors and describes the relationship between the rate of change of skin temperature and various parameters.

[0082] Together, these models form the human body's thermodynamic constraint system, providing a theoretical basis for subsequent energy optimization. Through these models, the system can accurately understand and predict the body's caloric needs, thereby optimizing energy allocation.

[0083] 3. Determine the energy optimization allocation strategy.

[0084] Based on the establishment of prediction model and constraint model, the system further determines the energy optimization allocation strategy.

[0085] like Figure 4As shown, the system uses a deep reinforcement learning algorithm to determine the heating mode and power supply based on predicted user motion and light data and a human body thermodynamic constraint model. The algorithm's state space includes the user's real-time motion pattern, predicted light intensity, wearable microgrid input power, and real-time skin temperature. The action space includes the maximum power supply time during the user's real-time motion state and the maximum power supply time of the energy harvesting layer. Preferably, the motion patterns in the state space can be categorized into multiple types, such as stationary, slow walking, fast walking, and running; light intensity is graded into three levels: low (0-200 lux), medium (200-1000 lux), and high (>1000 lux); and the wearable microgrid input power and skin temperature are precisely recorded.

[0086] like Figure 5 As shown, the system controls energy distribution among the wearable microgrid, supercapacitor, and flexible fiber battery based on the heating mode and the amount of heating power supplied, using a gradient game priority scheduling mechanism to achieve millisecond-level temperature regulation. This mechanism evaluates the priority of each energy unit based on energy availability, efficiency factor, user preference, and timeliness. Energy availability is typically expressed as the ratio of the current available energy to demand. The efficiency factor considers energy conversion and transmission efficiency. Generally, the conversion efficiency of photovoltaic units is between 10-18%, while the conversion efficiency of piezoelectric fibers is between 25-35%. User preferences are learned based on historical usage patterns. Timeliness indicates the urgency of meeting current needs and is typically expressed as a value between 0 and 1, with 0 indicating non-urgent and 1 indicating extreme urgency.

[0087] In practical applications, deep reinforcement learning algorithms continuously optimize energy allocation strategies through continuous interaction with the environment. For example, if the system detects that the user is in a low-temperature environment (e.g., outdoor temperature below 5°C) and is less active, it will increase heating power and select muscle heating mode. However, if the user is indoors and more active, it may reduce heating power or even disable heating mode to avoid energy waste.

[0088] In another embodiment of the present invention, Figure 2 As shown, the specific implementation of the dual-channel neural network further includes:

[0089] The motion prediction channel uses a single-layer long short-term memory neural network architecture with 128 hidden units and a variable time step between 0.3 and 300 seconds. This channel receives the user's motion pattern (e.g., still, walking, running) and time step data as input and outputs a sequence of the user's future motion patterns. Preferably, motion patterns can be categorized into 8-12 different types, including but not limited to still, slow walking, fast walking, running, jumping, and stair climbing.

[0090] The light prediction channel uses a multi-layer long short-term memory neural network architecture, consisting of an input layer, three hidden layers, and an output layer with 64, 128, and 64 hidden units, respectively. This channel receives current light intensity (in lux) and time data as input and outputs a sequence of future light intensities. In practice, light intensities range from 0 to 100,000 lux, which the system normalizes before inputting into the network.

[0091] In addition, the dual-channel network introduces an attention weight mechanism to improve prediction accuracy through adaptive weight adjustment. This mechanism is implemented using an improved Softmax function to assign dynamic weights to predictions at different time steps. For example, for the recent prediction (t+0.1 seconds), the initial weight coefficient Can be set to 0.6; for medium-term forecasts (t+5 minutes), weight compensation Can be set to 0.3; for long-term forecasts (t+30 minutes), weight compensation Can be set to 0.1. These weight coefficients satisfy and conditions, and can be dynamically adjusted according to the actual prediction results.

[0092] During the training phase, the system trains the network using historical data collected, including the motion patterns of different users in various environments and corresponding lighting conditions. Preferably, the training data consists of at least 1,000 hours of user activity records covering a variety of daily scenarios. This approach enables the dual-channel neural network to achieve over 95% short-term prediction accuracy and over 85% long-term prediction accuracy.

[0093] like Figure 3 As shown, the establishment of the human body thermodynamic constraint model of the present invention further includes:

[0094] The human thermal conductivity model describes the relationship between metabolic heat production and the skin's thermal conductivity, skin-to-air contact area, and temperature gradient. Under standard conditions (room temperature 22°C, 50% humidity), the thermal conductivity of human skin, λ, is typically 0.2-0.5 W / m·Kelvin, varying depending on individual differences. The skin-to-air contact area, A, is related to the user's body shape and is typically between 1.5 and 2.2 square meters. Temperature T refers to skin temperature, which is normally approximately 33-34°C. Length x represents the distance from the inside of the body to the skin surface, generally considered to be skin thickness, which is approximately 1-3 mm.

[0095] The human body heat dissipation model describes the relationship between the amount of heat exchange between the human body and the air, the skin surface heat dissipation coefficient, ambient temperature, and skin temperature. The skin surface heat dissipation coefficient α is affected by various factors, including air flow, humidity, and clothing insulation, and is typically between 4 and 6 watts / square meter·Kelvin. The ambient temperature Ta is the actual temperature of the user's environment, measured in real time by a temperature sensor. The skin temperature T is measured by a temperature sensor attached to the skin with an accuracy of ±0.1°C.

[0096] The human metabolic heat production model describes the relationship between metabolic heat production, metabolic power, and total heat production rate. Metabolic power Hm is closely related to the user's activity level, ranging from approximately 60-80 watts at rest, increasing to 100-120 watts during slow walking, and reaching 400-600 watts during strenuous exercise such as running. Total heat production rate R takes into account multiple factors, including basal metabolic rate, activity heat production, and additional heat production under special conditions (such as cold), and typically ranges from 0.8 to 1.2, with 1.0 representing normal activity.

[0097] The thermal balance model integrates the three aforementioned models to describe the relationship between the rate of change of skin temperature and various parameters. This model considers factors such as skin mass m0 (typically between 2 and 3 kg), skin heat capacity Cp (approximately 3.5-4.2 kJ / kg·Kelvin), metabolic heat, radiation, and heat exchange to calculate the rate of change of skin temperature dT / dt. Under steady-state conditions, dT / dt should be close to zero, indicating thermal balance. A positive dT / dt value indicates an increase in skin temperature, while a negative dT / dt value indicates a decrease in temperature.

[0098] Together, these four sub-models form a complete human body thermodynamic constraint system, providing a theoretical basis for optimal energy allocation. Through real-time monitoring and calculation, the system accurately assesses current thermal conditions and future thermal demands, thereby optimizing energy allocation strategies.

[0099] According to one embodiment of the present invention, the specific implementation of the heating mode further includes:

[0100] No-heating mode (Mode 0) is suitable for situations where the current body temperature has reached the target or the user's activity generates sufficient heat. The system determines whether heating is needed by monitoring the user's body temperature and activity status in real time. Preferably, this mode is selected when the skin temperature exceeds 33.5°C and the user is engaged in moderate or high-intensity activity (such as brisk walking or running). In this mode, the system does not supply power to the heating unit, using energy for other functions or storage. In no-heating mode, the system continues to monitor the user's status and is ready to switch to other modes when needed.

[0101] Skin Heating Mode (Mode 1) is suitable for situations where the ambient temperature is low but not extreme, requiring gentle heating. The system prefers this mode when the ambient temperature is between 10-20°C and the user is less active. In this mode, the heating element primarily acts on the surface of the skin, with a power setting between 1-3 watts / square centimeter. This heating method provides a quick increase in warmth while consuming relatively little energy.

[0102] Muscle Heating Mode (Mode 2) is suitable for low ambient temperatures and minimal user activity. This mode is typically selected when the ambient temperature drops to 0-10°C and the user is stationary or inactive. In this mode, the heating element targets deep muscle tissue, with a power setting typically between 2-4 watts / square centimeter. Muscle Heating Mode provides a longer-lasting warmth and is suitable for prolonged use in low-temperature environments.

[0103] Full Heating Mode (Mode 3) is designed for extremely cold environments or when the user requires special heating. The system selects this mode when the ambient temperature drops below 0°C or when the user explicitly requires maximum heating. In Full Heating Mode, the system heats both skin and muscle, reaching a power of 4-6 watts per square centimeter. This mode consumes the most energy but provides the strongest warmth, making it suitable for extreme conditions.

[0104] In actual applications, the selection of heating mode is based on a comprehensive combination of a real-time body temperature model, a heating mode classification model, and a heating capacity model. For example, if the system detects that the user's body temperature is more than 0.5°C below the normal value (36.4°C) and the ambient temperature is below 5°C, full heating mode will be prioritized. If the body temperature is only slightly below normal (no more than 0.3°C) and the ambient temperature is moderate, skin heating mode may be selected to save energy.

[0105] like Figure 4 As shown, the state space and action space design of the deep reinforcement learning algorithm of the present invention further includes:

[0106] The state space of the deep reinforcement learning algorithm includes four key elements: the user's real-time motion pattern, the predicted light intensity, the wearable microgrid input power, and the real-time skin temperature.

[0107] A user's real-time motion patterns are collected by piezoelectric fiber sensors and identified through a classification algorithm. These patterns are typically categorized as stationary (activity intensity <0.2, metabolic rate <1.2), slow walking (activity intensity 0.2-0.5, metabolic rate 1.2-2.0), fast walking (activity intensity 0.5-0.8, metabolic rate 2.0-3.0), and running (activity intensity >0.8, metabolic rate >3.0). This classification helps the system accurately estimate the user's metabolic heat production.

[0108] The predicted light intensity is obtained by the photovoltaic cells and light sensors and processed by a prediction algorithm. It is generally categorized into three light levels: low light (0-200 lux, typical indoor lighting), medium light (200-1000 lux, cloudy outdoors or bright indoors), and strong light (>1000 lux, sunny outdoors). Light intensity directly affects the efficiency of photovoltaic energy collection.

[0109] The wearable microgrid's input power represents the total energy currently available to the system, including energy harvested by the piezoelectric fibers, energy harvested by the photovoltaic cells, and energy stored in the supercapacitors and flexible fiber batteries. The system monitors this value in real time, expressed in milliwatt-hours (mWh), typically ranging from 0 to 500 mWh.

[0110] The real-time skin temperature is directly measured by a temperature sensor with an accuracy of ±0.1°C and a normal range of 32-35°C. This parameter is a key indicator for assessing user comfort and thermal needs.

[0111] The action space of the deep reinforcement learning algorithm includes two main parameters: the maximum power supply time of the user in real-time motion state and the maximum power supply time of the energy collection layer.

[0112] The maximum power supply time in the user's real-time motion state indicates the longest power supply time the system can maintain in the current motion state, in minutes, usually ranging from 30 to 480 minutes, depending on the current energy reserve and consumption rate.

[0113] The maximum power supply time of the energy harvesting layer indicates the maximum time the energy harvesting layer can support system operation under current environmental conditions. The unit is also minutes, ranging from 0 to infinity, and depends on environmental conditions (such as light intensity) and user activity status.

[0114] Within this state space and action space framework, the system guides the agent to learn the optimal strategy through a reward function. This reward function evaluates the difference between the predicted and measured values ​​and the difference between the real-time body temperature and the set temperature. Optimally, when the difference between the predicted and measured values ​​is less than 5% and the temperature deviation is less than 0.3°C, the system achieves a high reward; otherwise, it receives a low reward or a penalty. By continuously trying different strategies and evaluating the rewards, the system gradually learns the optimal energy allocation strategy.

[0115] In a preferred embodiment of the present invention, the deep reinforcement learning algorithm adopts an improved network structure to decompose the action-value function into a state-value part and an action-advantage part.

[0116] The state value component assesses the intrinsic value of the current state, that is, the value of the current state to the system without considering the specific action. For example, when a user is in a warm indoor environment and has moderate activity, the system is in a high state value because it is relatively easy to maintain the user's comfort regardless of the action taken. Conversely, if the user is in an extremely cold environment and has low energy reserves, the state value is lower.

[0117] The action advantage component assesses the relative advantage of performing a specific action in a specific state, that is, how good or bad the specific action is compared to the average action in that state. For example, selecting a full heating mode in a cold environment generally has a high action advantage; however, selecting the same mode during intense exercise may have a negative action advantage because the user has already generated sufficient heat.

[0118] This network architecture utilizes a shared feature extraction layer, independent value estimation streams, and independent action advantage estimation streams. The shared feature layer extracts key features from the input state; the value stream estimates the state value through two or three fully connected layers; and the advantage stream uses a similar structure to estimate the advantage of each action. Ultimately, the system integrates the state value and action advantage to form a complete Q-value estimate.

[0119] This structural design allows the system to more accurately distinguish between the value of the state itself and the impact of specific actions, reducing overestimation and improving learning efficiency. In practical applications, this decomposition method enables the system to achieve good performance within 50-100 training rounds, while traditional Q-networks may require 200-300 rounds. Furthermore, this structure helps improve the stability of the policy and reduce fluctuations during training, especially in scenarios with complex state transitions.

[0120] like Figure 5 As shown, the gradient game priority scheduling mechanism of the present invention further includes:

[0121] The gradient game priority scheduling mechanism first evaluates the priority of each energy unit based on energy availability, efficiency factor, user preference, and timeliness. Energy availability is calculated as the ratio of currently available energy to energy demand. A ratio greater than 1 indicates sufficient energy, while a ratio less than 1 indicates insufficient energy. The efficiency factor accounts for losses during energy conversion and transmission. For example, the conversion efficiency of photovoltaic cells can reach 15-20% under strong sunlight, but may drop to 10-12% under low light conditions. The conversion efficiency of piezoelectric fibers can reach 30-35% during intense exercise, but drops to 15-20% during light exercise. User preferences are derived by analyzing historical usage patterns. The system records user satisfaction feedback under different conditions and constructs a user preference model. Timeliness reflects the urgency of meeting current needs and is typically expressed as a value between 0 and 1. For example, if a user's body temperature is below 35.5°C, the timeliness may be assessed as above 0.9, indicating high urgency.

[0122] On this basis, each energy unit participates in resource allocation decisions as a game player. The system treats photovoltaic units, piezoelectric fiber units, and energy storage units as independent game entities. Each entity proposes a resource allocation plan based on its current state and forecast data. Through multiple rounds of iteration, these plans gradually converge to a Nash equilibrium point, forming the final energy allocation strategy.

[0123] The iterative process preferably uses a gradient descent method. Each participant adjusts its strategy in each iteration. The adjustment step size is initially set to 0.1 and then dynamically adjusted based on convergence. It can be increased to 0.2 when convergence is fast, and reduced to 0.05 or even lower when convergence is difficult. Typically, the system reaches a good equilibrium after 5-10 iterations.

[0124] After determining the final energy allocation plan at the equilibrium point, the system continues to optimize the strategy through a gradient update mechanism. This mechanism evaluates the performance of the current strategy based on actual results. For example, if the current strategy is found to be leading to excessive energy consumption or unstable temperature control, the strategy parameters are adjusted in the opposite direction. Preferably, the gradient direction is determined by both action advantage and state advantage, with a weight ratio of m:n = 0.6:0.4, which prioritizes the direct effects of actions. After multiple iterations, the system gradually converges to the optimal strategy, achieving dynamic coordination among the various components of the energy system.

[0125] In practice, this mechanism evaluates the current state and adjusts the strategy every 10-30 seconds to adapt to changes in the environment and user status. If the user's activity status changes dramatically (such as suddenly switching from sitting to running) or if environmental conditions change significantly (such as entering or leaving indoors), the system immediately triggers a reassessment to ensure the strategy remains optimal.

[0126] like Figure 6 As shown, the millisecond temperature regulation of the present invention further includes:

[0127] In a preferred embodiment of the present invention, the system achieves millisecond-level temperature regulation, specifically by performing 15 power switching cycles within 0.5 seconds, corresponding to a switching frequency of 30Hz. This high-frequency switching enables precise control of the heating module, avoiding the discomfort caused by drastic temperature fluctuations.

[0128] The system first determines the optimal heating mode based on the difference between the current heat demand and body temperature. For example, if the body temperature is at least 0.5°C below the normal value (36.4°C), the heat demand is assessed as "high"; if it is 0.2-0.5°C below, it is assessed as "medium"; and if it is below 0.2°C, it is assessed as "low". Taking into account the ambient temperature and user activity status, the system selects the most appropriate heating mode from four modes: no heating, skin heating, muscle heating, and full heating.

[0129] After determining the heating mode, the system calculates the required power level and determines the power switching parameters. The power level is dependent on the heat demand, heating mode, and environmental conditions. For example, in full heating mode, the power setting is 3 watts / square centimeter for low heat demand, 4 watts / square centimeter for medium heat demand, and 5-6 watts / square centimeter for high heat demand. The power switching parameter, PHeat, determines the charge and discharge states of the supercapacitor and flexible fiber battery. When PHeat is positive (e.g., +1), the supercapacitor and flexible fiber battery discharge; when it is negative (e.g., -1), the supercapacitor and flexible fiber battery charge.

[0130] The system preferably uses pulse-width modulation (PWM) to control power output, achieving fine-grained power control by adjusting the duty cycle. Under standard conditions, the duty cycle is adjustable from 10% to 90% in 5% increments. The power switching interval, τ, is typically set to 33.3 milliseconds (corresponding to a 30Hz switching frequency). Under special circumstances, it can be dynamically adjusted to a range of 20-50 milliseconds.

[0131] This millisecond-level temperature regulation mechanism enables the system to achieve smooth temperature transitions, avoiding the temperature fluctuations common in traditional systems. Actual tests have shown that this system can control temperature fluctuations within a range of ±0.5°C, significantly better than the ±2°C fluctuation range of traditional systems, greatly improving the user's comfort experience.

[0132] According to one embodiment of the present invention, the energy management of the flexible piezoelectric fiber and the flexible photovoltaic unit further comprises:

[0133] The system uses intelligent strategies to manage the energy harvesting and distribution of flexible piezoelectric fibers and flexible photovoltaic cells under varying environmental conditions. Specifically, when ambient light is sufficient (typically exceeding 1000 lux), the system prioritizes energy harvested by the flexible photovoltaic cells, as this provides a relatively stable energy input due to their high photoelectric conversion efficiency.

[0134] When the user's activity intensity is high (such as brisk walking or running, with a metabolic rate greater than 2.5), the system prioritizes energy harvested by flexible piezoelectric fibers because the piezoelectric effect is more pronounced at this time, improving energy harvesting efficiency. For example, during normal walking, each step generates approximately 0.1-0.3 millijoules of energy; while during running, this value increases to 0.3-0.8 millijoules.

[0135] When energy demand exceeds the currently harvested energy (e.g., a demand / supply ratio greater than 1.2), the system uses a supercapacitor array and flexible fiber batteries to provide supplemental energy. Supercapacitors typically have a discharge efficiency between 85-95%, enabling rapid response to changes in energy demand. The system prioritizes energy from supercapacitors, using the flexible fiber batteries only when the supercapacitors are insufficient, extending battery life.

[0136] Conversely, when the current harvested energy exceeds the energy demand (e.g., the demand / supply ratio is <0.8), the system stores the remaining energy in the supercapacitor array and flexible fiber battery. Preferably, the system maintains a 60-80% reserve in the energy storage unit to cope with sudden energy demands.

[0137] It's worth noting that energy allocation decisions aren't static; they're dynamically adjusted based on predicted future energy needs and energy harvesting conditions. The system leverages the aforementioned dual-channel neural network to predict user activity and environmental changes over the next 30 minutes, optimizing current energy allocation accordingly. For example, if it predicts a user is about to enter a cold outdoor environment from indoors, the system will increase energy reserves in advance. If it predicts a user is about to engage in strenuous exercise, the system might reduce heating power to accommodate the increased metabolic heat production.

[0138] This intelligent energy management strategy enables the system to maximize the use of ambient energy, reduce reliance on stored energy, and extend equipment operating time. Actual tests have shown that this dynamic adjustment mechanism can improve energy efficiency by 20-30% compared to a fixed strategy.

[0139] like Figure 7 , the implementation of the flexible piezoelectric fiber and the flexible photovoltaic unit of the present invention further includes:

[0140] Flexible piezoelectric fibers are made of flexible nano-piezoelectric materials. The material thickness is usually between 50-100 microns, and the length can be adjusted according to application requirements, usually between 5-15 cm. To enhance the piezoelectric effect, the material adopts a special nanostructure design to increase the surface area and deformation sensitivity. Nanoimprinting technology is used in the preparation process to form a regular nano-piezoelectric structure on a polyvinylidene fluoride (PVDF) substrate. The piezoelectric fiber is connected to the wearable microgrid through a flexible electrode. The electrode is made of silver nanowires or conductive polymer materials with a thickness controlled at 10-30 microns to ensure good conductivity and mechanical flexibility.

[0141] The flexible photovoltaic cell is fabricated using amorphous silicon using chemical vapor deposition. A transparent conductive material is printed onto a polyimide film, followed by plasma-enhanced chemical vapor deposition (PECVD) of an amorphous silicon layer (total thickness 200–500 nanometers). An aluminum or silver back electrode (approximately 200 nanometers thick) is then deposited by thermal evaporation. Finally, the flexible encapsulation is completed with ethylene-tetrafluoroethylene copolymer to enhance durability. The photovoltaic cell achieves a power conversion efficiency of 16% under standard testing conditions (1000 W / m² illumination and 25°C temperature).

[0142] The supercapacitor utilizes interdigitated electrodes made from MXene ink and an electrolyte made from a polyvinyl alcohol-based hydrogel. Using a high-resolution (50-100 micron) 3D printer, the MXene ink is used to form fine interdigitated electrodes. This structure significantly increases the effective surface area of ​​the electrodes and improves energy storage density. The electrode material, MXene, has a conductivity of 104 Siemens / cm and possesses high charge storage capacity, high rate capability, and hydrophilicity, far exceeding the performance of traditional electrode materials. The electrolyte, a polyvinyl alcohol-based hydrogel and lithium chloride salt solution, combined with the high conductivity of MXene, enables high-power, rapid charging and discharging, making it well-suited for flexible wearable devices.

[0143] Flexible fiber batteries use microfluidic wet spinning technology to prepare high-performance fiber batteries. First, MXene materials are mixed with LiMn2O4 (positive electrode) and Li4Ti5O 12 The (negative electrode) composite is oriented and arranged through microfluidic spinning to form a continuous fiber electrode (diameter 100-300μm), which significantly improves conductivity. Using a MXene-modified PEO-based gel electrolyte, the electrode fibers are coated by coaxial spinning to achieve an ultra-thin interface layer (<5μm) with an ionic conductivity of 10⁻³ S / cm. The positive and negative electrode fibers are woven into a spiral structure and coated with a self-healing polydimethylsiloxane encapsulation layer, achieving a capacity retention rate of >95% when the battery has a bending radius of <1mm.

[0144] The use of these flexible materials significantly improves the device's wearing comfort and energy harvesting efficiency. The flexible design allows the device to adapt to the body's curves and movements, reducing discomfort while maintaining excellent functionality. Actual testing shows that compared to rigid devices, this flexible design can improve user acceptance by 30-50% and extend the average daily wear time by 2-4 hours. In terms of energy harvesting, the optimized flexible material design increases piezoelectric energy harvesting efficiency by 25-35% and photovoltaic energy harvesting efficiency by 15-25%, significantly improving overall system performance.

[0145] In summary, this paper provides a wearable microgrid energy optimization and allocation method based on dynamic environmental perception. By organically combining technologies such as a multi-layer intelligent architecture, dual-channel prediction, a human body thermodynamic model, deep reinforcement learning decision-making, and gradient game scheduling, it achieves intelligent and refined energy management for wearable devices. This method offers significant advantages in energy efficiency, wearable comfort, and adaptability, providing new ideas and solutions for the development of wearable devices.

[0146] It should be noted that the above embodiments are only used to illustrate the present invention, and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A wearable microgrid energy optimization allocation method based on dynamic environment perception, characterized by: The following steps are involved: Obtain user motion data and ambient lighting data, including: The user's motion data is collected through flexible piezoelectric fibers set at the human joints; Ambient light data is collected through flexible photovoltaic units installed on the outer surface of clothing; Establish a dynamic environment prediction model and a human body thermodynamic constraint model, including: Based on the user motion data and the ambient light data, a dual-channel neural network is used to predict the user's motion pattern and light intensity change at the next moment, respectively, to obtain user motion prediction data and light prediction data; Establish a human body thermodynamic constraint model based on the thermal characteristics of the human body, including the human body thermal conductivity model, human body heat dissipation model, human body metabolic heat production model and thermal balance model; Determine energy optimization allocation strategies, including: Determining a heating mode and heating power using a deep reinforcement learning algorithm based on the user motion prediction data, the light prediction data, and the human body thermodynamic constraint model; According to the heating mode and the heating power, energy distribution control is performed on the wearable microgrid, supercapacitor and flexible fiber battery through a gradient game priority scheduling mechanism to achieve millisecond-level temperature regulation.

2. The wearable microgrid energy optimization allocation method based on dynamic environment perception according to claim 1 is characterized in that: The dual-channel neural network includes: Motion prediction channel, used to receive user motion pattern and time step data and output the user's future motion pattern sequence; Light prediction channel, which receives current light intensity and time data and outputs future light intensity sequence; Among them, the motion prediction channel and the illumination prediction channel both adopt a long short-term memory neural network structure, and introduce an attention weight mechanism to improve prediction accuracy.

3. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: The establishment of the human body thermodynamic constraint model includes: Human thermal conductivity model: describes the relationship between human metabolic heat production and skin thermal conductivity, skin-air contact area, and temperature gradient; Human body heat dissipation model: describes the relationship between the heat exchange between the human body and the air, the skin surface heat release coefficient, the ambient temperature and the skin temperature; Human metabolic heat production model: describes the relationship between human metabolic heat production, metabolic power and total heat production rate; Thermal balance model: describes the relationship between the rate of change of skin temperature and skin mass, heat ratio, metabolic heat, radiation and heat exchange between the human body and the air.

4. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: The heating modes include: No heating mode: Select this mode when the current body temperature has reached the standard or the user's activities generate enough heat; Skin heating mode: Select when the ambient temperature is low but not extreme and slight heating is required; Muscle heating mode: Select this mode when the ambient temperature is low and the user's activity level is low, and deep heating is required; Full heating mode: Select this mode when the environment is extremely cold or the user has special needs and needs to heat both the skin and muscles; Among them, the selection of heating mode is determined based on the real-time body temperature model, heating mode classification model and heating amount model.

5. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: The state space of the deep reinforcement learning algorithm includes: the user's real-time motion pattern; the predicted light intensity; the wearable microgrid input power; the real-time skin temperature; the action space of the deep reinforcement learning algorithm includes: the maximum power supply time under the user's real-time motion state; the maximum power supply time of the energy collection layer; The system is guided to learn the optimal strategy through a profit function, which evaluates the difference between the predicted value and the measured value and the temperature difference between the real-time body temperature and the set temperature value.

6. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 5 is characterized in that: The deep reinforcement learning algorithm adopts an improved network structure to decompose the action value function into a state value part and an action advantage part, where: The state value component assesses the intrinsic value of the current state; The action advantage component assesses the relative advantage of taking a specific action in a specific state; Improve learning efficiency and policy stability by estimating state value and action advantage separately.

7. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: The gradient game priority scheduling mechanism includes: Evaluate the priority of each energy unit based on energy availability, efficiency factor, user preference and timeliness; Treat each energy unit as a game participant and reach the equilibrium point of resource allocation through iterative calculation; Evaluate the current strategy based on actual results and adjust the allocation parameters along the gradient direction; After multiple iterations, the optimal energy allocation strategy is converged to achieve dynamic coordination among the components of the energy system.

8. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: The millisecond-level temperature adjustment specifically includes: Execute 15 power switches in 0.5 seconds to achieve fine control of the heating module; Determine the optimal heating mode based on the current heat demand and body temperature difference; Calculate required power levels and determine power switching parameters; The charge and discharge states of the supercapacitor are controlled according to the power switching parameter. When the power parameter is positive, the supercapacitor and the flexible fiber battery are discharged. When the power parameter is negative, the supercapacitor and the flexible fiber battery are charged.

9. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: Energy management of the flexible piezoelectric fiber and the flexible photovoltaic unit includes: When the ambient light is sufficient, the energy collected by the flexible photovoltaic unit is used first; When the user's activity intensity is high, the energy harvested by the flexible piezoelectric fiber is preferentially used; When the energy demand is greater than the current harvested energy, supplementary energy is provided by supercapacitor arrays and flexible fiber batteries; When the current harvested energy is greater than the energy demand, the remaining energy is stored in the supercapacitor array and flexible fiber battery; Among them, energy allocation decisions are dynamically adjusted based on predicted future energy demand and energy collection conditions.

10. The wearable microgrid energy optimization distribution method based on dynamic environment perception according to claim 1 is characterized in that: The implementation of the flexible piezoelectric fiber and the flexible photovoltaic unit includes: The flexible piezoelectric fiber is made of flexible nano-piezoelectric material and is connected to the wearable microgrid via flexible electrodes; The flexible photovoltaic unit is prepared by plasma enhanced chemical vapor method using amorphous silicon material on a polyimide substrate, and finally flexibly encapsulated with ethylene-tetrafluoroethylene copolymer; The supercapacitor uses an interdigitated structure made of MXene ink and an electrolyte made of polyvinyl alcohol-based hydrogel, and the 3D interdigitated structure is formed by 3D printing technology; The flexible fiber battery combines MXene materials with LiMn2O4 (positive electrode) and Li4Ti5O 12 (Anode) composite, prepared by microfluidic wet spinning technology; Among them, the application of the flexible material improves the wearing comfort and energy collection efficiency of the device.