Methods for controlling seat firmness, training methods and equipment for parameter generation models
Patent Information
- Application Number
- CN202610929128.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]但是,上述方式无法自适应地调整座椅的软硬程度,缺乏灵活性
[0009]本公开实施例提供的技术方案与现有技术相比具有如下优点:
Smart Images

Figure CN122724367A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of seat control technology, and in particular to a method for controlling seat firmness, a method for training a parameter generation model, and a device. Background Technology
[0002] With the development of intelligent automotive technology, consumers have increasingly higher requirements for driving and riding comfort, and car seats are one of the important components that determine driving and riding comfort.
[0003] Currently, improvements in car seat comfort largely rely on material or structural optimization. For example, by using rigid foam materials or comfort sponges during design or modification, seat comfort can be improved. Alternatively, the seat posture can be changed according to the user's operation, such as reclining the seat back.
[0004] However, the above methods cannot adaptively adjust the firmness of the seat, lacking flexibility. Summary of the Invention
[0005] To address the aforementioned technical issues, this disclosure provides a method for controlling seat firmness, a method for training a parameter generation model, and an apparatus.
[0006] A first aspect of this disclosure provides a method for controlling the firmness of a seat, including: Obtain the status information corresponding to the target seat, the status information including the vehicle's driving status information and / or the pressure sensing information of the target seat, and the target seat is provided with at least one support component; Based on the status information, determine the target control parameters; The support strength of the at least one support component is controlled according to the target control parameters.
[0007] A second aspect of this disclosure provides a method for training a parameter generation model, comprising: Acquire at least one sample data point corresponding to the target seat. The sample data includes the sample status information of the target seat and the sample control parameters corresponding to the sample status information. The sample status information includes sample driving status information and / or sample pressure sensing information. For each piece of sample data, a reward value corresponding to the sample data is determined based on the sample status information. The reward value is used to characterize the degree of matching between the sample status information and the sample control parameters. The initial model is trained based on the at least one sample data and the reward value to obtain a parameter generation model.
[0008] A third aspect of this disclosure provides an electronic device, including: processor; Memory, used to store executable instructions; The processor is used to read executable instructions from memory and execute the executable instructions to implement the methods provided in the first or second aspect above.
[0009] The technical solution provided in this disclosure has the following advantages compared with the prior art: The seat firmness control method, parameter generation model training method, and device provided in this disclosure dynamically adjust the control parameters of the support components on the seat according to the vehicle driving state and pressure sensing data, thereby changing the support force provided by the support components to the user. This enables the user to have a more suitable riding experience under the current driving state, realizing real-time and dynamic control of the seat firmness and improving flexibility. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a method for controlling the firmness of a seat according to an embodiment of this disclosure; Figure 2 A schematic diagram of a target seat provided in an embodiment of this disclosure; Figure 3 A flowchart illustrating a training method for a parameter generation model provided in this embodiment of the disclosure; Figure 4 This is a schematic diagram of preset pressure distribution information provided in an embodiment of the present disclosure; Figure 5 A flowchart of sample data acquisition provided for embodiments of this disclosure; Figure 6 A flowchart illustrating a method for controlling seat firmness according to another embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of a seat firmness control device provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of the structure of a training device for a parameter generation model provided in an embodiment of this disclosure; Figure 9This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0013] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0014] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0015] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0016] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] Figure 1 This is a flowchart illustrating a method for controlling seat firmness according to an embodiment of this disclosure. This method can be executed by a seat firmness control device, which can be implemented using software and / or hardware. The seat firmness control device can be configured in an electronic device, such as a server or terminal, where the terminal specifically includes a vehicle infotainment system, mobile phone, computer, or tablet computer. It is understood that the seat firmness control method provided in this disclosure can also be applied to other scenarios.
[0019] The following is about Figure 1 The method for controlling the firmness of the seat, as shown, will be introduced. Figure 1 As shown, the method for controlling the firmness of a seat provided in this embodiment includes the following steps.
[0020] S101. Obtain the status information corresponding to the target seat.
[0021] The target seat refers to the car seat that currently requires adjustment of firmness, including the driver's seat, the front passenger seat, or other seats.
[0022] Status information is used to reflect the current working environment and usage status of the seat. Status information includes vehicle driving status information and / or pressure sensing information of the target seat.
[0023] Vehicle driving status information refers to various parameters that characterize the vehicle's current driving conditions and reflect its dynamic driving status in real time, including whether the vehicle is in motion, its speed, the magnitude of acceleration, and changes in its direction of travel. These parameters are acquired through various sensors mounted on the vehicle.
[0024] The pressure sensing information of the target seat shows the pressure distribution between the user and the seat, reflecting the user's body characteristics, sitting posture, and the current support effect of the seat on the user.
[0025] The target seat is equipped with at least one support component. The support component can change the support strength of the seat surface by changing its own state. Each target seat is equipped with at least one support component, or multiple support components can be arranged according to the structural zones of the seat to achieve differentiated adjustment of the softness and hardness of different areas of the seat.
[0026] Electronic devices collect vehicle driving status information and seat pressure sensor information. When the vehicle is stationary and there is no change in driving status, the seat firmness can be adjusted based solely on the pressure sensor information. When the vehicle is in motion, the seat firmness is adjusted by combining the vehicle driving status information and the seat pressure sensor information.
[0027] S102. Determine the target control parameters based on the status information.
[0028] The target control parameters are used to control the working state of the support components, thereby changing the support strength of the target seat. The value of the target control parameters is determined by the currently acquired state information; different state information corresponds to different target control parameters to achieve adaptive adjustment of the seat's firmness.
[0029] S103. Control the support strength of at least one support component according to the target control parameters.
[0030] Support strength refers to the magnitude of the supporting force exerted by the supporting components on the surface of the target seat, which determines the firmness of the target seat. The greater the support strength, the greater the supporting force on the surface of the target seat, and the firmer the target seat; the smaller the support strength, the smaller the supporting force on the surface of the target seat, and the softer the target seat.
[0031] The target control parameters include the control parameters of each support component. The electronic device sends the control command corresponding to the control parameter to the drive device of the support component. The drive device adjusts the working state of the support component according to the command, thereby changing the support strength.
[0032] This embodiment dynamically adjusts the control parameters of the support components on the seat according to the vehicle's driving status and pressure sensor data, changing the support force provided by the support components to the user, so that the seat's firmness matches the current driving status or different users' body shapes, achieving adaptive adjustment of the seat's firmness and improving flexibility and comfort.
[0033] Figure 2 This is a schematic diagram of a target seat provided as an embodiment of the present disclosure. Figure 2 As shown, the supporting component is an air bag. The target control parameters include at least one of the following: target inflation rate, target inflation time, and target pressure holding time.
[0034] For each air bag, the support strength of at least one support component is controlled according to target control parameters, including: controlling the air bag to be inflated at a target inflation rate; and / or controlling the air bag to be inflated until the inflation time reaches the target inflation time; and / or controlling the amount of air in the air bag to remain constant during the time period corresponding to the target pressure holding time.
[0035] The target inflation rate refers to the volume of gas inflated into a single airbag per unit time. Based on the target inflation rate, the electronic system sends corresponding control commands to the air pump, adjusting operating parameters such as the pump's output power and valve opening to stabilize the volume of gas inflated into the airbag at the target inflation rate. By controlling the inflation speed, a smooth transition during seat firmness adjustment is achieved, avoiding any impact or jerking sensation during adjustment that could affect the user's riding experience.
[0036] The target inflation time refers to the duration from when the air pump starts inflating the air bag until inflation stops. Electronic equipment determines the final inflation volume of the air bag by controlling the inflation time, thereby determining the final support strength of the seat and enabling precise adjustment of the seat's firmness. This improves the responsiveness and reliability of seat firmness adjustment.
[0037] The target pressure holding time refers to the duration during which the internal air pressure and support strength of the airbag remain stable after it reaches the preset target inflation volume. After the airbag reaches the target inflation volume and stops inflating, the electronic system maintains the inflation volume and corresponding support strength within the target pressure holding time, effectively ensuring the continuous stability of the seat's support strength.
[0038] Optionally, the airbag can be combined with a pressure sensor to form a pressure-sensitive airbag. The pressure sensor is positioned on top of the airbag to maximize its contact with the user, improving the accuracy of pressure sensing information acquisition and preventing inaccurate pressure distribution information due to insufficient contact between the user and the seat during airbag inflation. Therefore, the pressure sensing information of the target seat includes the pressure sensing information corresponding to each support component.
[0039] For example, the target seat includes a seat surface and a backrest, and at least one support member is arranged on the seat surface and the backrest, respectively. The support height of the support member is positively correlated with the distribution distance; and / or, the arrangement density of the support members is negatively correlated with the distribution distance. The distribution distance is the distance between the center of the support member and the center of the seat surface or backrest where the support member is located.
[0040] In other words, the airbag arrangement is determined based on the specific shape of the target seat, so that the airbags form a wrapping shape relative to the user, providing a sense of envelopment. At the same time, because the pressure varies as it extends outward from the spine, the airbags are arranged with denser airbags on the inside and looser airbags on the outside, further improving the comfort of the target seat.
[0041] In some embodiments, the target control parameters are determined based on the state information and the parameter generation model. That is, the state information is input into the parameter generation model to obtain the target control parameters output by the parameter generation model.
[0042] The training method for the parameter generation model will be introduced below. Figure 3 This is a flowchart illustrating a training method for a parameter generation model provided in an embodiment of this disclosure. Figure 3 As shown, the method includes the following steps.
[0043] S301. Obtain at least one sample data point corresponding to the target seat.
[0044] S302. For each sample data, determine the reward value corresponding to the sample data based on the sample status information.
[0045] S303. Train the initial model based on at least one sample data and the reward value to obtain the parameter generation model.
[0046] The sample data includes the sample status information of the target seat and the sample control parameters corresponding to the sample status information.
[0047] Among them, sample status information refers to the information on seat usage status and vehicle driving conditions collected during the training process, including sample driving status information and / or sample pressure sensing information.
[0048] Sample control parameters refer to the control parameters of the support components of the target seat under the corresponding sample state information.
[0049] In some embodiments, an evaluation network is used to participate in the training process of the parameter generation model. Specifically, with the reward value corresponding to the output sample data as the objective, a first preset model is trained based on at least one sample data to obtain an evaluation model; with the objective of maximizing the evaluation value output by the evaluation model, a second preset model is trained based on the sample state information in at least one sample data to obtain a policy generation model.
[0050] The second preset model is used to output initial control parameters based on the sample state information, and the evaluation model is used to output evaluation values based on the initial control parameters. The evaluation values are used to characterize the degree of matching between the sample state information and the initial control parameters.
[0051] The first preset model is used to learn the mapping relationship between sample data and reward values. It takes sample state information and sample control parameters as input and outputs the corresponding evaluation value. By continuously adjusting the parameters of the first preset model, the evaluation value output by the first preset model is made as close as possible to the reward value pre-labeled for each sample data.
[0052] The second preset model is used to learn the mapping relationship between sample state information and optimal control parameters. It takes sample state information as input and outputs the corresponding control parameters. By continuously adjusting the parameters of the second preset model, the evaluation model is made to obtain the largest possible evaluation value when the initial control parameters output by the second preset model are input into it.
[0053] This disclosure improves the accuracy and objectivity of value estimation by training an evaluation model, providing a reliable foundation for training the parameter generation model. Furthermore, by using the evaluation model to guide the training of the parameter generation model, the efficiency and strategy of parameter optimization are improved, while enabling the model to continuously self-optimize.
[0054] The reward value is used to characterize the degree of matching between sample state information and sample control parameters, in order to quantitatively evaluate the quality of the sample control parameters under the corresponding sample state information. Its value range is usually a continuous numerical range, and the higher the value, the better the effect of the corresponding sample control parameter, that is, the more well the sample control parameter matches the sample state information.
[0055] Specifically, the reward value corresponding to the sample data is determined based on the pressure difference information and the sample driving status information.
[0056] The sample pressure sensing information includes sample pressure distribution information, and the pressure difference information is used to characterize the difference between the sample pressure distribution information and the preset pressure distribution information.
[0057] The sample pressure distribution information is collected by multiple pressure detection units arranged inside the target seat, reflecting the pressure magnitude and distribution pattern of each area of the user's contact surface with the seat.
[0058] Figure 4 This is a schematic diagram of preset pressure distribution information provided in an embodiment of the present disclosure, such as... Figure 4 As shown, the preset pressure distribution information refers to the pressure distribution standard calibrated based on ergonomic principles and a large amount of real vehicle comfort test data, which can provide the best riding experience for most users.
[0059] Pressure difference information is used to quantify the degree of difference between the sample pressure distribution and the preset pressure distribution, including multiple dimensions such as pressure magnitude difference and pressure distribution pattern difference. The smaller the difference between the sample pressure distribution and the preset pressure distribution, the closer the current pressure distribution is to the ideal state, and the better the seat support effect.
[0060] This embodiment of the disclosure standardizes and objectifies the evaluation of seat support effect by introducing preset pressure distribution information as a reference standard, ensuring the consistency and reliability of the evaluation results. Furthermore, the preset pressure distribution information can be flexibly adjusted according to different seat designs and different user groups, enabling the model to adapt to diverse product needs and user preferences.
[0061] In some embodiments, the sample driving status information includes vehicle speed and speed duration, and the sample pressure sensing information includes user presence time. Different reward value calculation methods are used according to different vehicle speeds.
[0062] When the vehicle speed is greater than or equal to a preset speed threshold, the reward value of the sample data is determined based on pressure difference information, speed duration, and sample driving status information.
[0063] The speed duration is the duration during which the vehicle's speed is greater than or equal to a preset speed threshold.
[0064] A preset speed threshold is used to distinguish between high-speed and medium-to-low-speed driving states. Sample driving state information also includes vehicle acceleration, lateral acceleration, steering wheel angle, braking signal, etc. These parameters can reflect the dynamic driving conditions of the vehicle, such as acceleration, deceleration, and turning.
[0065] During high-speed driving, which is typically accompanied by longer driving times, the cumulative effect of user fatigue is more pronounced. Therefore, the calculation of reward values in high-speed scenarios requires the integration of multi-dimensional information: pressure difference information is used to assess the basic support effect of the seat, speed duration is used to assess the degree of fatigue accumulation, and other sample driving state information is used to assess the impact of the vehicle's dynamic operating conditions on support requirements.
[0066] When the vehicle speed is less than a preset speed threshold but greater than zero, the reward value of the sample data is determined based on the pressure difference information and the sample driving status information.
[0067] At low to medium speeds, the vehicle's dynamic operating conditions are relatively simple, and users' needs for seats focus more on a balance between comfort and support. Therefore, the reward value calculation in low to medium speed scenarios mainly combines pressure difference information and other sample driving state information, prioritizing the improvement of ride comfort while ensuring basic support.
[0068] When the vehicle speed is zero, the reward value of the sample data is determined based on the pressure difference information and the user's time in place.
[0069] In a stationary state, the impact of dynamic operating parameters such as acceleration and steering angle on seat support does not need to be considered. At this time, the user's primary needs are a comfortable riding experience and fatigue relief after long periods of sitting. Therefore, the reward value calculation in a stationary scenario mainly combines pressure difference information and the user's time in position, prioritizing riding comfort and dynamically adjusting the seat firmness based on the duration of sitting to alleviate fatigue.
[0070] This embodiment divides the vehicle's driving state into three different scenarios: high speed, medium-low speed, and stationary. It also designs a targeted reward value calculation logic for each scenario, enabling the model to learn the optimal control strategy under different scenarios and achieve adaptive adjustment of seat firmness in all scenarios.
[0071] In some embodiments, when the vehicle speed is greater than or equal to a preset speed threshold, the reward value of the sample data is determined based on pressure difference information, speed duration, and sample driving state information, including: determining a first index value based on speed duration, wherein speed duration is negatively correlated with the first index value; and determining the reward value of the sample data based on pressure difference information, the first index value, and sample driving state information.
[0072] The first indicator value is used to quantify the cumulative fatigue effect in high-speed driving scenarios. In the initial stages of high-speed driving, when the speed duration is short, the first indicator value is high, and the reward value calculation primarily focuses on seat support. As driving time increases and the speed duration gradually lengthens, the first indicator value gradually decreases. At this point, the reward value calculation gradually increases the requirements for comfort, alleviating driving fatigue over long periods. This dynamic adjustment mechanism can effectively improve the comfort of long-distance driving while ensuring driving safety.
[0073] For example, the first index value can be the product of the first encouragement parameter and the speed duration. When the high speed duration is long and the seat becomes too hard, the first encouragement parameter is reduced in time.
[0074] In some embodiments, when the vehicle speed is zero, determining the reward value of the sample data based on pressure difference information and user presence time includes: determining a second indicator value based on user presence time, wherein user presence time is negatively correlated with the second indicator value; and determining the reward value of the sample data based on pressure difference information, the second indicator value, and sample driving status information.
[0075] The second indicator value is used to quantify the cumulative fatigue effect under static conditions. In the initial stage of static seating, when the user's time in the seat is short, the second indicator value is relatively high, and the reward value calculation primarily focuses on the seat's support. As the user's time in the seat gradually increases, the second indicator value gradually decreases, and the reward value calculation gradually increases the requirements for comfort, enhancing the comfort of long-term static seating.
[0076] For example, the second indicator value could be the product of a second incentive parameter and the user's on-duty time. When the vehicle is stationary, the second incentive parameter is promptly reduced if the on-duty time is too long.
[0077] This embodiment of the disclosure achieves a dynamic balance between support and comfort through a first index value and a second index value, enabling the reward value to dynamically adjust and optimize the target as the riding time increases. This allows the control strategy learned by the model to more accurately match changes in user needs, further improving the flexibility of seat firmness control.
[0078] Optionally, when the vehicle speed is greater than or equal to a preset speed threshold, the reward value is calculated as follows: R[Total Reward] = 0.2 Pressure difference information +0.2 Overall vehicle acceleration +0.1 Lateral acceleration +0.1 The absolute value of the difference between the current steering wheel angle and the previous steering wheel angle, plus 0.2. (Brake signal) (Check value) + 0.1 (Expert review value reward) (Check value) + 0.1 (Speed duration) First encouragement parameter).
[0079] When the vehicle speed is less than a preset speed threshold but greater than zero, the reward value is calculated as follows: R[Total Reward] = 0.3 Pressure difference information +0.2 Overall vehicle acceleration +0.1 Lateral acceleration +0.1 The absolute value of the difference between the current steering wheel angle and the previous steering wheel angle, plus 0.2. (Brake signal) (Check value) + 0.1 (Expert review value reward) (Check value).
[0080] When the vehicle speed is zero, the reward value is calculated as follows: R[Total Reward] = 0.4 Pressure difference information +0.3 (Expert review value reward) (Check value) + 0.3 (User's time in office) Second encouragement parameter).
[0081] The expert review value bonus refers to the bonus value parameter obtained based on expert experience. The check value is used to increase the bonus value of the braking system and avoid the bonus being diluted due to the signal value being too small.
[0082] Understandably, the weights of the above parameters can be calibrated on a real vehicle as needed.
[0083] Figure 5 This is a flowchart illustrating a sample data acquisition method provided in an embodiment of this disclosure. Figure 5 As shown, the sample data collection process includes the following steps: S3011. During the target sampling period, acquire the first state information and the first control parameters corresponding to the target seat, and write them into the preset buffer.
[0084] The duration of the sampling period is related to the vehicle's communication cycle. The shorter the communication cycle, the shorter the sampling period, and the more samples are collected.
[0085] The preset buffer adopts a first-in-first-out storage mechanism to store sample status information and sample control parameters in the most recent sampling periods in chronological order, including the first status information and the first control parameters.
[0086] Optionally, the collected data can be differentiated based on whether the vehicle speed is zero, which facilitates differential training during subsequent reinforcement learning and avoids abnormal comfort caused by seat movement during driving.
[0087] Optionally, determine whether the sample pressure sensing information is valid. If the pressure in the sample pressure sensing information is greater than a preset pressure threshold, or if an occupancy signal of the target seat is detected, it is determined that there is a real occupant in the target seat, and the sample pressure sensing information is valid. Otherwise, it is determined that there is no real occupant in the target seat, or that there is pressure generated by an object placed on the target seat or by squeezing the target seat, and the sample pressure sensing information is determined to be invalid, and the current first state information and first control parameters are discarded.
[0088] Optionally, after confirming the validity of the sample pressure sensing information, the sample driving status information and the sample pressure sensing information are fitted into a univariate array and written into a preset buffer.
[0089] For example, the first state information is fitted into a unary array. The structure of the unary array is: [vehicle speed, braking signal, acceleration signal, longitudinal acceleration signal, steering wheel angle signal, backrest pressure sensing signal (i), seat cushion pressure sensing signal (k)].
[0090] S3012. If, among at least one sample data that has been collected, there exists a sample data whose sample status information is consistent with the first status information but whose sample control parameters are inconsistent with the first control parameters, then a sample data is determined based on the second status information and the first control parameters in the preset buffer.
[0091] The second state information is the state information preceding the first state information.
[0092] Alternatively, if there is no sample data whose sample status information matches the first status information among at least one collected sample data, then a sample data is determined based on the first status information and the first control parameter.
[0093] Because the status information and control parameters come from different modules, their generation, transmission, and processing times differ, as do the times they are written to the preset buffer. Identical status information should correspond to identical control parameters. If the same status corresponds to different control parameters, it indicates a mismatch during sample data acquisition; the current first control parameter actually corresponds to the status information of the previous sampling period. Therefore, the correct pairing should be the second status information and the first control parameter, defining the second status information and the first control parameter as a single sample data point.
[0094] Furthermore, the first state information and the first control parameters are saved.
[0095] This embodiment of the disclosure mitigates the asynchronous transmission of state information and control parameters by using a preset buffer, identifies mismatched data by verifying data uniqueness, and converts mismatched data into valid sample data by correcting the mismatched data, thereby avoiding data waste and ensuring the quality and utilization rate of sample data.
[0096] Figure 6 A flowchart illustrating a method for controlling seat firmness according to another embodiment of this disclosure. Figure 6 As shown, the method includes the following steps: S301. Obtain at least one sample data point corresponding to the target seat.
[0097] S302. Determine the reward value corresponding to each sample data.
[0098] S3031. With the reward value corresponding to the output sample data as the objective, train the first preset model based on at least one sample data to obtain the evaluation model.
[0099] S3032. With the goal of maximizing the evaluation value output by the evaluation model, the second preset model is trained based on the sample state information in at least one sample data to obtain the policy generation model.
[0100] S101. Obtain the status information corresponding to the target seat.
[0101] S102. Input the state information into the parameter generation model to obtain the target control parameters output by the parameter generation model.
[0102] S103. Control the support strength of at least one support component according to the target control parameters.
[0103] Understandably, in the actual target seat adjustment process, the state information used and the generated target control parameters can be used as sample data to continuously optimize and train the evaluation model and the parameter generation model.
[0104] Figure 7 This is a schematic diagram of a seat firmness control device provided in an embodiment of this disclosure.
[0105] In this embodiment of the disclosure, the seat firmness control device can be located within an electronic device, and can be understood as a part of the functional modules of the aforementioned electronic device. For example... Figure 7As shown, the seat firmness control device 70 may include a first acquisition module 71, a first determination module 72, and a control module 73; wherein, the first acquisition module 71 is used to acquire the status information corresponding to the target seat, the status information including the vehicle's driving status information and / or the pressure sensing information of the target seat, and at least one support component is arranged on the target seat; the first determination module 72 is used to determine the target control parameters based on the status information; the control module 73 is used to control the support strength of at least one support component based on the target control parameters.
[0106] In some embodiments, the first determining module 72 is used to input state information into the parameter generation model to obtain the target control parameters output by the parameter generation model; wherein, the parameter generation model is trained in the following manner: acquiring at least one sample data corresponding to the target seat, the sample data including sample state information of the target seat and sample control parameters corresponding to the sample state information, the sample state information including sample driving state information and / or sample pressure sensing information; for each sample data, determining the reward value corresponding to the sample data according to the sample state information, the reward value being used to characterize the degree of matching between the sample state information and the sample control parameters; training the initial model based on at least one sample data and the reward value to obtain the parameter generation model.
[0107] In some embodiments, the sample pressure sensing information includes sample pressure distribution information, and determining the reward value corresponding to the sample data based on the sample state information includes: Based on the pressure difference information and the sample driving status information, the reward value corresponding to the sample data is determined. The pressure difference information is used to characterize the difference between the sample pressure distribution information and the preset pressure distribution information.
[0108] In some embodiments, the sample driving state information includes vehicle speed and speed duration, and the sample pressure sensing information includes user presence time. Determining the reward value corresponding to the sample data based on the sample state information includes: when the vehicle speed is greater than or equal to a preset speed threshold, determining the reward value of the sample data based on pressure difference information, speed duration, and sample driving state information, where the speed duration is the duration during which the vehicle speed is greater than or equal to the preset speed threshold; when the vehicle speed is less than the preset speed threshold but greater than zero, determining the reward value of the sample data based on pressure difference information and sample driving state information; and when the vehicle speed is zero, determining the reward value of the sample data based on pressure difference information and user presence time.
[0109] In some embodiments, determining the reward value of sample data based on pressure difference information, speed duration, and sample driving state information includes: determining a first indicator value based on speed duration, wherein speed duration is negatively correlated with the first indicator value; and determining the reward value of sample data based on pressure difference information, the first indicator value, and sample driving state information.
[0110] In some embodiments, determining the reward value of sample data based on pressure difference information and user on-site time includes: determining a second indicator value based on user on-site time, wherein user on-site time is negatively correlated with the second indicator value; and determining the reward value of sample data based on pressure difference information, the second indicator value, and sample driving status information.
[0111] In some embodiments, acquiring at least one sample data corresponding to the target seat includes: acquiring first state information and first control parameters corresponding to the target seat during the target sampling period and writing them into a preset buffer; if there is sample data whose sample state information is consistent with the first state information and whose sample control parameters are inconsistent with the first control parameters among the at least one sample data already acquired, then a sample data is determined based on the second state information and the first control parameters in the preset buffer, wherein the second state information is the previous state information of the first state information.
[0112] In some embodiments, training an initial model based on at least one sample data point and a reward value to obtain a parameter generation model includes: training a first preset model based on at least one sample data point with the reward value corresponding to the output sample data as the objective, to obtain an evaluation model; and training a second preset model based on sample state information in at least one sample data point with the objective of maximizing the evaluation value output by the evaluation model, to obtain a policy generation model; wherein the second preset model is used to output initial control parameters according to the sample state information, the evaluation model is used to output an evaluation value according to the initial control parameters, and the evaluation value is used to characterize the degree of matching between the sample state information and the initial control parameters.
[0113] In some embodiments, the support component is an air bag, and the target control parameters include at least one of a target inflation rate, a target inflation time, and a target holding time; the control module 73 is used to control the air bag to be inflated according to the target inflation rate; and / or to control the air bag to be inflated until the inflation time reaches the target inflation time; and / or to control the amount of air in the air bag to remain constant during the time period corresponding to the target holding time.
[0114] It should be noted that, Figure 7 The seat firmness control device 70 shown can perform the various steps in the above method embodiments and achieve the various processes and effects in the above method embodiments, which will not be elaborated here.
[0115] Figure 8 This is a schematic diagram of the structure of a training device for a parameter generation model provided in an embodiment of this disclosure. Figure 8 As shown, the training device 80 for the parameter generation model includes a second acquisition module 81, a second determination module 82, and a training module 83. The second acquisition module 81 is used to acquire at least one sample data corresponding to the target seat. The sample data includes sample state information of the target seat and sample control parameters corresponding to the sample state information. The sample state information includes sample driving state information and / or sample pressure sensing information. The second determination module 82 is used to determine the reward value corresponding to each sample data based on the sample state information. The reward value is used to characterize the degree of matching between the sample state information and the sample control parameters. The training module 83 is used to train the initial model based on at least one sample data and the reward value to obtain the parameter generation model.
[0116] In some embodiments, the sample pressure sensing information includes sample pressure distribution information. The second determining module 82 is used to determine the reward value corresponding to the sample data based on the pressure difference information and the sample driving status information. The pressure difference information is used to characterize the difference between the sample pressure distribution information and the preset pressure distribution information.
[0117] In some embodiments, the sample driving state information includes vehicle speed and speed duration, and the sample pressure sensing information includes user presence time. The second determining module 82 includes a first determining unit, a second determining unit, and a third determining unit. The first determining unit is used to determine the reward value of the sample data based on pressure difference information, speed duration, and sample driving state information when the vehicle speed is greater than or equal to a preset speed threshold, wherein the speed duration is the duration during which the vehicle speed is greater than or equal to the preset speed threshold; the second determining unit is used to determine the reward value of the sample data based on pressure difference information and sample driving state information when the vehicle speed is less than the preset speed threshold but greater than zero; the third determining unit is used to determine the reward value of the sample data based on pressure difference information and user presence time when the vehicle speed is zero.
[0118] It should be noted that, Figure 8 The training device 80 for the parameter generation model shown can execute the various steps in the above method embodiments and realize the various processes and effects in the above method embodiments, which will not be elaborated here.
[0119] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0120] like Figure 9 As shown, the electronic device may include a processor 910 and a memory 920 storing computer program instructions.
[0121] Specifically, the processor 910 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this disclosure.
[0122] Memory 920 may include a large-capacity storage for information or instructions. For example, and not limitingly, memory 920 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 920 may include removable or non-removable (or fixed) media. Where appropriate, memory 920 may be internal or external to the integrated gateway device. In a particular embodiment, memory 920 is a non-volatile solid-state memory. In a particular embodiment, memory 920 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0123] The processor 910 reads and executes computer program instructions stored in the memory 920 to perform the steps of the seat firmness control method and / or parameter generation model training method provided in the embodiments of this disclosure.
[0124] In one example, the electronic device may also include a transceiver 930 and a bus 940. Wherein, as... Figure 9 As shown, the processor 910, memory 920 and transceiver 930 are connected via bus 940 and communicate with each other.
[0125] Bus 940 may include hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 940 may include one or more buses.
[0126] This disclosure also provides a computer-readable storage medium that can store a computer program. When the computer program is executed by a processor, the processor enables the processor to implement the seat firmness control method and / or parameter generation model training method provided in this disclosure.
[0127] The aforementioned storage medium may, for example, include a memory 920 containing computer program instructions, which can be executed by a processor 910 of an electronic device to complete the seat firmness control method and / or parameter generation model training method provided in the embodiments of this disclosure. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), compact disc ROM (CD-ROM), magnetic tape, floppy disk, and optical data storage device.
[0128] This disclosure also provides a vehicle that includes electronic devices that can implement the various processes and effects described in the above embodiments of this disclosure, which will not be elaborated here.
[0129] This disclosure also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, they implement the seat firmness control method and / or parameter generation model training method provided in this disclosure, and can achieve the various processes and effects in the above embodiments of this disclosure, which will not be elaborated here.
[0130] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for controlling the firmness of a seat, characterized in that, include: S101. Obtain the status information corresponding to the target seat, the status information including the vehicle's driving status information and / or the pressure sensing information of the target seat, and at least one support component is arranged on the target seat; S102. Determine the target control parameters based on the status information; S103. Control the support strength of the at least one support component according to the target control parameters.
2. The method according to claim 1, characterized in that, Determining the target control parameters based on the status information includes: The state information is input into the parameter generation model to obtain the target control parameters output by the parameter generation model; The parameter generation model is trained in the following manner: S301. Obtain at least one sample data corresponding to the target seat. The sample data includes sample status information of the target seat and sample control parameters corresponding to the sample status information. The sample status information includes sample driving status information and / or sample pressure sensing information. S302. For each piece of sample data, determine the reward value corresponding to the sample data based on the sample status information. The reward value is used to characterize the degree of matching between the sample status information and the sample control parameters. S303. The initial model is trained based on the at least one sample data and the reward value to obtain a parameter generation model.
3. The method according to claim 2, characterized in that, The sample pressure sensing information includes sample pressure distribution information, and determining the reward value corresponding to the sample data based on the sample state information includes: Based on the pressure difference information and the sample driving status information, the reward value corresponding to the sample data is determined. The pressure difference information is used to characterize the difference between the sample pressure distribution information and the preset pressure distribution information.
4. The method according to claim 3, characterized in that, The sample driving status information includes vehicle speed and speed duration, the sample pressure sensing information includes user presence time, and determining the reward value corresponding to the sample data based on the sample status information includes: When the vehicle speed is greater than or equal to a preset speed threshold, the reward value of the sample data is determined based on the pressure difference information, the speed duration, and the sample driving state information, wherein the speed duration is the duration during which the vehicle speed is greater than or equal to the preset speed threshold. When the vehicle speed is less than the preset speed threshold but greater than zero, the reward value of the sample data is determined based on the pressure difference information and the sample driving status information. When the vehicle speed is zero, the reward value of the sample data is determined based on the pressure difference information and the user's time in place.
5. The method according to claim 4, characterized in that, The step of determining the reward value of the sample data based on the pressure difference information, the speed duration, and the sample driving state information includes: A first index value is determined based on the speed duration, wherein the speed duration is negatively correlated with the first index value; The reward value of the sample data is determined based on the pressure difference information, the first indicator value, and the sample driving status information.
6. The method according to claim 4, characterized in that, Determining the reward value of the sample data based on the pressure difference information and the user's time in place includes: A second indicator value is determined based on the user's in-service time, and the user's in-service time is negatively correlated with the second indicator value; The reward value of the sample data is determined based on the pressure difference information, the second indicator value, and the sample driving status information.
7. The method according to any one of claims 2-6, characterized in that, The step of obtaining at least one sample data point corresponding to the target seat includes: S3011. During the target sampling period, acquire the first state information and the first control parameters corresponding to the target seat, and write them into a preset buffer. S3013. In at least one sample data that has been collected, if there is a sample data whose sample status information is consistent with the first status information, but whose sample control parameters are inconsistent with the first control parameters, then a sample data is determined based on the second status information and the first control parameters in the preset buffer, wherein the second status information is the previous status information of the first status information.
8. The method according to any one of claims 2-6, characterized in that, The step of training the initial model based on the at least one sample data and the reward value to obtain a parameter generation model includes: S3031. With the goal of outputting the reward value corresponding to the sample data, train the first preset model based on the at least one sample data to obtain the evaluation model; S3032. With the goal of maximizing the evaluation value output by the evaluation model, a second preset model is trained based on the sample state information in the at least one sample data to obtain a policy generation model; wherein, the second preset model is used to output initial control parameters according to the sample state information, the evaluation model is used to output an evaluation value according to the initial control parameters, and the evaluation value is used to characterize the degree of matching between the sample state information and the initial control parameters.
9. The method according to claim 1, characterized in that, The supporting component is an air bag, and the target control parameters include at least one of the target inflation rate, target inflation time, and target pressure holding time. For each of the air bags, controlling the at least one support component according to the target control parameters includes: Control the inflation of the airbag at a target inflation rate; and / or, Control the inflation of the air bag until the inflation time reaches the target inflation time; and / or, During the time period corresponding to the target pressure holding time, the inflation volume in the air bag is kept constant.
10. A training method for a parameter generation model, characterized in that, The method includes: S301. Obtain at least one sample data corresponding to the target seat. The sample data includes the sample status information of the target seat and the sample control parameters corresponding to the sample status information. The sample status information includes sample driving status information and / or sample pressure sensing information. S302. For each piece of sample data, determine the reward value corresponding to the sample data based on the sample status information. The reward value is used to characterize the degree of matching between the sample status information and the sample control parameters. S303. The initial model is trained based on the at least one sample data and the reward value to obtain a parameter generation model.
11. The method according to claim 10, characterized in that, The sample pressure sensing information includes sample pressure distribution information, and determining the reward value corresponding to the sample data based on the sample state information includes: Based on the pressure difference information and the sample driving status information, the reward value corresponding to the sample data is determined. The pressure difference information is used to characterize the difference between the sample pressure distribution information and the preset pressure distribution information.
12. The method according to claim 11, characterized in that, The sample driving status information includes vehicle speed and speed duration; the sample pressure sensing information includes user presence time; and determining the reward value corresponding to the sample data based on the sample status information includes: When the vehicle speed is greater than or equal to a preset speed threshold, the reward value of the sample data is determined based on the pressure difference information, the speed duration, and the sample driving state information, wherein the speed duration is the duration during which the vehicle speed is greater than or equal to the preset speed threshold. When the vehicle speed is less than the preset speed threshold but greater than zero, the reward value of the sample data is determined based on the pressure difference information and the sample driving status information. When the vehicle speed is zero, the reward value of the sample data is determined based on the pressure difference information and the user's time in place.
13. An electronic device, characterized in that, include: processor; Memory, used to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method of any one of claims 1-12.