Electric toy car remote speed control system and method based on scene recognition

By installing sensors on electric toy cars to create a multimodal twin scene model, and combining it with user operating habits to generate driving decisions, the stability and personalization problems of traditional electric toy cars in complex scenarios are solved, and intelligent speed and route adjustment are achieved.

CN121500831APending Publication Date: 2026-02-10HEBEI WANFENGZHIHUANG BICYCLE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511592522.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional electric toy cars cannot automatically adjust their speed and route according to different scenarios, leading to collisions and jamming issues, and failing to provide a personalized user driving experience.

Method used

By installing multiple sensors to collect real-time multimodal data, a multimodal twin scene model is established. Initial driving decisions are generated by combining user operation habit factors, and the model is automatically updated when the user uploads instructions.

Benefits of technology

It enables electric toy cars to drive stably in complex scenarios, avoiding collisions and jamming, and improving driving efficiency and the personalization of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121500831A_ABST
    Figure CN121500831A_ABST
Patent Text Reader

Abstract

The invention discloses an electric toy car remote speed control system and method based on scene recognition, and relates to the technical field of toy car control. Various sensors are installed on the electric toy car, then real-time multi-modal data frames of a scene where the electric toy car is located are collected, a multi-modal twin scene model is established according to the real-time multi-modal data frames, historical user operation records are obtained, and then user operation habit factors in different driving areas are obtained according to the historical user operation records. The user operation habit factors are bound to all the driving available areas, then an initial driving decision is generated and executed according to the scene object distribution in the driving available areas and the user operation habit factors, and when a user uploads an operation instruction, the initial driving decision is automatically updated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of toy car control, in particular to a remote speed control system and method for electric toy cars based on scene recognition. BACKGROUND

[0002] As an entertainment product deeply loved by children and toy enthusiasts, the function and intelligence level of electric toy cars have always been an important direction for the development of the industry. The traditional speed control method of electric toy cars is relatively single, usually only realizing basic operations such as simple forward, backward, acceleration, and deceleration, lacking effective perception of the driving scene and in-depth analysis of user operation habits.

[0003] In actual use, electric toy cars face complex and variable driving scenes, such as indoor furniture layout and outdoor terrain and topography. However, the existing control system cannot automatically adjust the speed and driving route according to the characteristics of different scenes, leading to problems such as collision and jamming during driving, reducing the user's experience.

[0004] In addition, different users have different habits and preferences when operating electric toy cars, for example, some users prefer fast driving, while some users prefer slow and stable driving. The traditional control system does not fully consider these individual differences and cannot provide personalized driving experience for users. SUMMARY

[0005] The purpose of the present application is to provide a remote speed control system and method for electric toy cars based on scene recognition to solve the problems in the background art.

[0006] In order to achieve the above-mentioned purpose, the present application provides the following technical solutions: A remote speed control method for electric toy cars based on scene recognition, comprising the following steps: Step S1, installing multiple sensors on the electric toy car to collect real-time multi-modal data frames of the scene, and establishing a multi-modal twin scene model according to the real-time multi-modal data frames; Step S2, obtaining historical user operation records, and then obtaining user operation habit factors under different drivable areas according to the historical user operation records; Step S3, binding the user operation habit factors to each drivable area, and then generating an initial driving decision according to the scene object distribution in the drivable area and the user operation habit factors, and executing the initial driving decision, and automatically updating the initial driving decision when the user uploads an operation instruction.

[0007] Further, the process of real-time scene data includes: The sensor collects multiple real-time scene data of the user operation scene with a fixed data collection period, sets a collection timestamp for all sensor data, corrects real-time scene data with inconsistent sampling intervals by using a linear interpolation method, and then completes time alignment between each real-time scene data; A mapping relationship between the sensor coordinate system and the scene world coordinate system is established, a local space coordinate system is established with the rear wheel shaft center of the electric toy car as the origin, the forward direction as the X axis, and the vertical ground upward as the Z axis; For optical cameras, infrared sensors and other sensors, Zhang Zhengyou calibration method is used, and the detection results of all sensors are mapped to the local space coordinate system; According to the time alignment results of each real-time scene data, each real-time scene data in the local space coordinate system is associated and compressed by frame, and then real-time multi-modal data frames are obtained.

[0008] Further, the process of establishing a multi-modal twin scene model according to the real-time multi-modal data frame includes: For any real-time multi-modal data frame, the data frame of the corresponding video data is converted into point cloud data, the extent of the scene and the reflection characteristics of the obstacles are judged based on the sound wave signal data frame, and the scene object semantics of each point cloud data is given according to the reflection characteristics of the obstacles, and then a single frame of point cloud corresponding to the real-time multi-modal data frame is obtained; The single frame point clouds of consecutive frames are spliced, the point clouds in the overlapping area are fused by using the mean value method, and the scene object semantics of each point is preserved, and then a multi-modal twin scene model is obtained; The points in the multi-modal twin scene model that are adjacent and have the same scene object semantics are connected, and a plurality of scene feature semantic regions are divided in the multi-modal twin scene model according to the connection; The scene feature semantic region includes a drivable region, an obstacle and a boundary region, wherein the drivable region is labeled with characteristic physical properties according to the corresponding scene type; A toy car model is set in the center position of the multi-modal twin scene model in proportion, and the acceleration, angular velocity data, real-time driving speed and other data under the corresponding frame are labeled on the toy car model, and the multi-modal twin scene model is automatically updated every data collection period.

[0009] Further, the basic operation information includes user ID, operation timestamp, operation instruction type, operation instruction parameter, operation number and operation instruction timestamp; The scene association information includes the historical multi-modal twin scene model and the semantic labeling result of the drivable region during operation; The driving state information includes the real-time speed, acceleration, driving direction and distance to the nearest obstacle of the electric toy car at each driving timestamp during operation.

[0010] Further, the process of obtaining the user operation habit factor under different drivable areas according to the historical user operation records comprises: According to the spatial distribution of the drivable areas in the historical multi-modal twin scene model, the historical user operation records are divided into a plurality of historical operation records, each of which contains the basic operation information, scene association information, and driving state information of the electric toy car during driving in different drivable areas. Further, the user operation habit factor under different drivable areas is obtained for the corresponding user according to the historical operation records, and the user operation habit factor includes the operation frequency factor f a , the habit speed factor f b , and the operation response factor f c . Each time a new historical user operation record is uploaded by the user, the user operation habit factor is automatically updated.

[0011] Further, the process of binding the user operation habit factor to each drivable area comprises: According to the drivable area corresponding to the user operation habit factor, the drivable area in the multi-modal twin scene model is matched, and the user operation habit factor is bound to the corresponding drivable area according to the matching result. If it is determined that there is a drivable area in the multi-modal twin scene model that is not bound to any user operation habit factor, the physical properties of the drivable area are matched with the drivable area corresponding to the user operation habit factor, the user operation habit factor corresponding to the drivable area with the closest physical property value is selected, and is recorded as a temporary user operation habit factor, and is bound to the drivable area that is not bound to any user operation habit factor.

[0012] Further, the process of generating an initial driving decision comprises: If the user uploads a driving endpoint through the remote control software, the current position of the electric toy car is taken as the starting coordinate, and the scene feature semantic area with obstacle and boundary area mark is taken as the priority condition to generate a plurality of estimated driving routes. Each estimated driving route is mapped to the multi-modal twin scene model, and the drivable area type through which the estimated driving route passes is labeled, and then the actual cost g(n) and the estimated cost h(n) of each estimated driving route are obtained according to the drivable area with the user operation habit factor through which each estimated driving route passes. The estimated driving route is divided into driving nodes with equal length, and one or more operation instructions are bound to each driving node according to the nearest scene feature semantic area of the driving node. The estimated driving route with the smallest sum of actual cost g(n) and estimated cost h(n) is selected as the initial driving decision and sent to the electric toy car and remote control software. The electric toy car then drives according to the initial driving decision and automatically executes the operation instructions bound to each driving node. If the user sends an operation command to the toy electric car through remote control software and has set the destination, the estimated driving route will be regenerated and executed based on the position coordinates of the toy electric car when the operation command ends, until the toy electric car reaches the destination. If the user sends an operation command to the toy electric car through remote control software without setting a destination, the operation command will be executed automatically and no estimated route will be generated.

[0013] A remote speed control system for electric toy cars based on scene recognition includes a scene perception module, a habit analysis module, and a driving control module. The scene perception module is used to collect real-time scene data of the scene, perform spatiotemporal alignment of various real-time scene data to generate corresponding real-time multimodal data frames, and establish a multimodal twin scene model based on the real-time multimodal data frames. The habit analysis module is used to obtain historical user operation records, and then obtain user operation habit factors under different drivable areas based on the historical user operation records. The driving control module is used to bind user operation habit factors to each driving area, and then generate an estimated driving route based on the distribution of scene objects in the driving area and user operation habit factors. It obtains the actual cost and estimated cost of each estimated driving route, and then selects the estimated driving route with the minimum actual cost and estimated cost to generate an initial driving decision and execute it. When the user uploads operation instructions, the initial driving decision is automatically updated.

[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention collects real-time scene data and generates multimodal data frames through spatiotemporal alignment, thereby establishing a multimodal twin scene model. This enables electric toy cars to accurately perceive the characteristics of the scene they are in, and automatically adjust their speed and driving route according to different scenes, effectively avoiding collisions and jams, and improving the driving stability and safety of toy cars in complex scenes.

[0015] 2. This invention, by comprehensively considering both actual and estimated costs when generating the estimated driving route, and selecting the route with the lowest cost as the initial driving decision, enables the electric toy car to plan routes more efficiently and rationally during driving, thereby improving driving efficiency. At the same time, when the user uploads operation instructions, the initial driving decision is automatically updated, realizing the organic combination of user intervention and intelligent decision-making, and enhancing the flexibility and adaptability of the system. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0017] Figure 1 This is a flowchart of a method for remote speed control of an electric toy car based on scene recognition, as described in this invention.

[0018] Figure 2 This is a system framework diagram of a scene recognition-based remote speed control system for electric toy cars according to the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figure 1 As shown, a remote speed control method for an electric toy car based on scene recognition includes the following steps: Step S1: Install various sensors on the electric toy car to collect real-time multimodal data frames of the scene, and establish a multimodal twin scene model based on the real-time multimodal data frames; Step S2: Obtain historical user operation records, and then obtain user operation habit factors under different drivable areas based on the historical user operation records; Step S3: Bind user operation habit factors to each drivable area, and then generate and execute initial driving decisions based on the distribution of scene objects in the drivable area and user operation habit factors. When the user uploads operation commands, the initial driving decisions are automatically updated.

[0021] Step S1 is achieved through the following process: Step S101: Collect real-time scene data. The specific process includes: The electric toy car is equipped with various types of sensors, including but not limited to optical cameras, infrared sensors, acoustic sensors, wheeled odometers, and lidar. The sensor collects multiple real-time scene data of the user's operation scenario at a fixed data acquisition cycle, including optical video data, acoustic signal data, infrared video data, depth image data, acceleration and angular velocity data, etc., and the duration of the data acquisition cycle is generally 10ms to 50ms. Since all real-time scene data are collected centered on electric toy cars, the data collection objects for all real-time scene data are the same. Set acquisition timestamps for all sensor data with an accuracy of 1ms. Use linear interpolation to correct real-time scene data with inconsistent sampling intervals (e.g., infrared video data is 66.7ms / frame, which needs to be aligned with optical video data of 33.3ms / frame, so insert one frame of interpolated data between two frames of infrared data) to complete the time alignment between various real-time scene data. Establish a mapping relationship between the sensor coordinate system and the scene world coordinate system. Take the rear wheel axle of the electric toy car as the origin, the forward direction as the X-axis, and the vertical upward direction as the Z-axis to establish a local spatial coordinate system. For sensors such as optical cameras and infrared sensors, the Zhang Zhengyou calibration method is used, along with acoustic sensors (based on sound sources at known locations, correcting the spatial position deviation of the array), to map the detection results of all sensors to a local spatial coordinate system; Based on the time alignment results of various real-time scene data, the various real-time scene data in the local spatial coordinate system are frame-by-frame associated and compressed to obtain real-time multimodal data frames.

[0022] Step S102: Establish a multimodal twin scene model based on real-time multimodal data frames. The specific process includes: For any real-time multimodal data frame, the corresponding video data frame is converted into point cloud data. The degree of openness and obstacle reflection characteristics in the scene are determined based on the sound wave signal data frame. The scene object semantics (such as walls, furniture, floors, etc.) are assigned to each point cloud data according to the obstacle reflection characteristics, thereby obtaining the single frame point cloud corresponding to the real-time multimodal data frame. The point clouds of consecutive frames are stitched together, and the point clouds of overlapping areas (points with a distance of <0.5cm are considered to be overlapping) are fused using the mean method, while preserving the scene object semantics of each point, thus obtaining a multimodal twin scene model; Connect adjacent points in the multimodal twin scene model that have the same scene object semantics, and divide the multimodal twin scene model into several scene feature semantic regions based on the connections. The scene feature semantic region includes drivable areas (such as floors, carpets, and sand), obstacles (such as furniture, toys, and walls), and boundary areas (such as table edges and stairwells). The drivable areas are marked with characteristic physical properties, including friction coefficient and smoothness, depending on the corresponding scenario. Since the real-time scene data corresponding to the multimodal twin scene model is collected centered on an electric toy car, a toy car model is set up proportionally at the center of the multimodal twin scene model, and the acceleration, angular velocity data, real-time driving speed and other data of the corresponding frame are marked on the toy car model. The multimodal twin scene model is automatically updated after each data collection cycle.

[0023] Step S2 is achieved through the following process: Step S201: Obtain historical user operation records. The specific process includes: Historical user operation records mainly come from users' control behavior of electric toy cars through remote control software. Users register user accounts when using the remote control software, and then whenever users use the remote control software, the remote control software automatically collects user operation records, scene association information, and driving status information. The basic operation information includes user ID, operation timestamp, toy car device ID, operation command type (accelerate, decelerate, turn left, turn right, stop), operation command parameters (acceleration / deceleration range, etc.), number of operations, and timestamp of each operation command. The scene association information includes the historical multimodal twin scene model during operation and the semantic annotation results of the drivable area (scene object type, obstacle distribution, etc.), wherein the generation process of the historical multimodal twin scene model is the same as step S1; The driving status information includes the real-time speed, acceleration, driving direction, and distance to the nearest obstacle of the electric toy car at each driving time stamp when the operation is performed.

[0024] Step S202: Obtain user operation habit factors for different drivable areas based on historical user operation records. The specific process includes: Based on the spatial distribution of drivable areas in the historical multimodal twin scene model, the historical user operation records are divided into multiple historical operation records. Each historical operation record contains basic operation information, scene association information, and driving status information of the electric toy car during its driving in different drivable areas. Then, based on historical operation records, the corresponding user is obtained, and the user operation habit factor is determined under different drivable areas. The user operation habit factor includes an operation frequency factor f. a Habitual speed factor f b and the operation response factor f c ; The operation frequency factor is obtained by statistically analyzing the percentage of each type of operation command used within the drivable area: ,in This represents the operation frequency factor for the k-th operation instruction. and The number of occurrences of the k-th and i-th operation commands is represented by I, which represents the total number of operation command types and is a natural number greater than 4. The habitual speed factor is a statistical measure of the range of driving speeds commonly used by users in this scenario area. The speed data is fitted using a normal distribution, and the mean value μ is taken. v As a benchmark habitual speed, the standard deviation σ v As a measure of speed fluctuation range: , Where M represents the total number of operations. This represents the real-time speed during the j-th operation; the operation response factor is obtained by statistically analyzing the user's operation response time to changes in the drivable scene, i.e., the time interval from when the sensor detects a scene change (such as a decrease in obstacle distance) to when the user issues an operation command. ,in The timestamp of the q-th operation instruction. The scene change detection timestamp associated with the qth operation instruction (obtained through the timestamp of the toy car model passing through the drivable area in the historical multimodal twin model); the user operation habit factor is automatically updated whenever the user uploads a new historical user operation record.

[0025] Step S3 is achieved through the following process: Step S301: Bind user operation habit factors to each drivable driving area. The specific process includes: Based on step S2, generate the drivable area corresponding to the user operation habit factor, match the drivable area in the multimodal twin scene model, and bind the user operation habit factor to the corresponding drivable area according to the matching result. If it is determined that there is a drivable area in the multimodal twin scene model that is not bound to any user operation habit factors, then the physical attributes of the drivable area are matched with the drivable area corresponding to the user operation habit factor. The user operation habit factor corresponding to the drivable area with the closest physical attribute value is selected and recorded as a temporary user operation habit factor, and bound to the drivable area that is not bound to any user operation habit factor.

[0026] Step S302: Generate initial driving decision, the specific process of which includes: If the user uploads the destination via remote control software, the current position of the electric toy car will be used as the starting coordinates, and several estimated driving routes will be generated with the priority of avoiding scene feature semantic regions with obstacles and boundary area markers. Each predicted driving route is mapped onto a multimodal twin scene model, and the types of drivable areas traversed by the predicted driving routes are labeled. Then, based on the user operation habit factors in the drivable areas traversed by each predicted driving route, the actual cost g(n) and the predicted cost h(n) of each predicted driving route are obtained. The estimated driving route is divided into driving nodes with equal spacing. Based on the nearest scene feature semantic region of the driving node, one or more operation instructions are bound to each driving node. For example, if there is a scene feature semantic region with obstacles near the driving node, then the driving node is set to move to the right and decelerate according to the direction of the estimated driving route. ; in and Let num represent the actual cost and total length of the nth estimated route, respectively. n This represents the total number of travel nodes for the nth estimated travel route. The parameters were obtained through several experiments to correct them; ,in , , , These represent the planar coordinates of the first and last travel nodes of the nth estimated travel route, respectively. The estimated driving route with the smallest sum of actual cost g(n) and estimated cost h(n) is selected as the initial driving decision and sent to the electric toy car and remote control software. The electric toy car then drives according to the initial driving decision and automatically executes the operation instructions bound to each driving node. If the user sends an operation command to the toy electric car through remote control software and has set the destination, the estimated driving route will be regenerated and executed based on the position coordinates of the toy electric car when the operation command ends, until the toy electric car reaches the destination. If the user sends an operation command to the toy electric car through remote control software without setting a destination, the operation command will be executed automatically and no estimated route will be generated.

[0027] Please see Figure 2 As shown, a remote speed control system for electric toy cars based on scene recognition includes a scene perception module, a habit analysis module, and a driving control module. The scene perception module is used to collect real-time scene data of the scene, perform spatiotemporal alignment of various real-time scene data to generate corresponding real-time multimodal data frames, and establish a multimodal twin scene model based on the real-time multimodal data frames. The habit analysis module is used to obtain historical user operation records, and then obtain user operation habit factors under different drivable areas based on the historical user operation records. The driving control module is used to bind user operation habit factors to each driving area, and then generate an estimated driving route based on the distribution of scene objects in the driving area and user operation habit factors. It obtains the actual cost and estimated cost of each estimated driving route, and then selects the estimated driving route with the minimum actual cost and estimated cost to generate an initial driving decision and execute it. When the user uploads operation instructions, the initial driving decision is automatically updated.

[0028] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for remote speed control of an electric toy car based on scene recognition, characterized in that, Includes the following steps: Step S1: Install various sensors on the electric toy car to collect real-time multimodal data frames of the scene, and establish a multimodal twin scene model based on the real-time multimodal data frames; Step S2: Obtain historical user operation records, and then obtain user operation habit factors under different drivable areas based on the historical user operation records; Step S3: Bind user operation habit factors to each drivable area, and then generate and execute initial driving decisions based on the distribution of scene objects in the drivable area and user operation habit factors. When the user uploads operation commands, the initial driving decisions are automatically updated.

2. The method for remote speed control of an electric toy car based on scene recognition according to claim 1, characterized in that, The process of real-time scene data includes: The sensor collects multiple real-time scene data of the user operation scenario at a fixed data acquisition cycle, sets a collection timestamp for all sensor data, and uses linear interpolation to correct real-time scene data with inconsistent sampling intervals, thereby completing the time alignment between various real-time scene data. Establish a mapping relationship between the sensor coordinate system and the scene world coordinate system. Take the rear wheel axle of the electric toy car as the origin, the forward direction as the X-axis, and the vertical upward direction as the Z-axis to establish a local spatial coordinate system. For sensors such as optical cameras and infrared sensors, the Zhang Zhengyou calibration method is used, along with acoustic sensors, to map the detection results of all sensors to a local spatial coordinate system; Based on the time alignment results of various real-time scene data, the various real-time scene data in the local spatial coordinate system are frame-by-frame associated and compressed to obtain real-time multimodal data frames.

3. The method for remote speed control of an electric toy car based on scene recognition according to claim 2, characterized in that, The process of building a multimodal twin scene model based on real-time multimodal data frames includes: For any real-time multimodal data frame, the corresponding video data frame is converted into point cloud data. The degree of openness and obstacle reflection characteristics in the scene are judged based on the sound wave signal data frame. The scene object semantics are assigned to each point cloud data according to the obstacle reflection characteristics, thereby obtaining the single-frame point cloud corresponding to the real-time multimodal data frame. The point clouds of consecutive frames are stitched together, and the point clouds of overlapping areas are fused using the mean method, while preserving the scene object semantics of each point, thus obtaining a multimodal twin scene model. Connect adjacent points in the multimodal twin scene model that have the same scene object semantics, and divide the multimodal twin scene model into several scene feature semantic regions based on the connections. The scene feature semantic region includes drivable areas, obstacles, and boundary areas, wherein the drivable areas are labeled with characteristic physical attributes according to the corresponding scene type; A toy car model is set up at the center of the multimodal twin scene model according to scale, and the acceleration, angular velocity data, real-time driving speed and other data of the corresponding frame are marked on the toy car model. The multimodal twin scene model is automatically updated after each data collection cycle.

4. The method for remote speed control of an electric toy car based on scene recognition according to claim 3, characterized in that, The historical user operation records include user operation records, scene association information, and driving status information; The basic operation information includes user ID, operation timestamp, operation instruction type, operation instruction parameters, number of operations, and timestamp of each operation instruction. The scene association information includes the historical multimodal twin scene model during operation and the semantic annotation results of the drivable area; The driving status information includes the real-time speed, acceleration, driving direction, and distance to the nearest obstacle of the electric toy car at each driving time stamp when the operation is performed.

5. The method for remote speed control of an electric toy car based on scene recognition according to claim 4, characterized in that, The process of obtaining user operation habit factors for different drivable areas based on historical user operation records includes: Based on the spatial distribution of drivable areas in the historical multimodal twin scene model, the historical user operation records are divided into multiple historical operation records. Each historical operation record contains basic operation information, scene association information, and driving status information of the electric toy car during its driving in different drivable areas. Then, based on historical operation records, the corresponding user is obtained, and the user operation habit factor is determined under different drivable areas. The user operation habit factor includes an operation frequency factor f. a Habitual speed factor f b and the operation response factor f c Whenever a user uploads a new historical user operation record, the user operation habit factor is automatically updated.

6. The method for remote speed control of an electric toy car based on scene recognition according to claim 5, characterized in that, The process of linking user operating habits to various drivable driving areas includes: Based on the drivable area corresponding to the user operation habit factor, match the drivable area in the multimodal twin scene model, and bind the user operation habit factor to the corresponding drivable area according to the matching result; If it is determined that there is a drivable area in the multimodal twin scene model that is not bound to any user operation habit factors, then the physical attributes of the drivable area are matched with the drivable area corresponding to the user operation habit factor. The user operation habit factor corresponding to the drivable area with the closest physical attribute value is selected and recorded as a temporary user operation habit factor, and bound to the drivable area that is not bound to any user operation habit factor.

7. The method for remote speed control of an electric toy car based on scene recognition according to claim 6, characterized in that, The process of generating the initial driving decision includes: If the user uploads the destination via remote control software, the current position of the electric toy car will be used as the starting coordinates, and several estimated driving routes will be generated with the priority of avoiding scene feature semantic regions with obstacles and boundary area markers. Each estimated driving route is mapped onto a multimodal twin scene model, and the types of drivable areas traversed by the estimated driving routes are labeled. Then, based on the user operation habit factors in the drivable areas traversed by each estimated driving route, the actual cost and estimated cost of each estimated driving route are obtained. The estimated driving route is divided into driving nodes with equal spacing, and one or more operation instructions are bound to each driving node based on the semantic region of the scene feature closest to the driving node. The estimated driving route with the minimum sum of actual cost and estimated cost is selected as the initial driving decision and sent to the electric toy car and remote control software. The electric toy car then drives according to the initial driving decision and automatically executes the operation instructions bound to each driving node. If the user sends an operation command to the toy electric car through remote control software and has set the destination, the estimated driving route will be regenerated and executed based on the position coordinates of the toy electric car when the operation command ends, until the toy electric car reaches the destination. If the user sends an operation command to the toy electric car through remote control software without setting a destination, the operation command will be executed automatically and no estimated route will be generated.

8. A scene-recognition-based remote speed control system for electric toy cars, used to implement the scene-recognition-based remote speed control method for electric toy cars as described in any one of claims 1-7, characterized in that, It includes a scene perception module, a habit analysis module, and a driving control module; The scene perception module is used to collect real-time scene data of the scene, perform spatiotemporal alignment of various real-time scene data to generate corresponding real-time multimodal data frames, and establish a multimodal twin scene model based on the real-time multimodal data frames. The habit analysis module is used to obtain historical user operation records, and then obtain user operation habit factors under different drivable areas based on the historical user operation records. The driving control module is used to bind user operation habit factors to each driving area, and then generate an estimated driving route based on the distribution of scene objects in the driving area and user operation habit factors. It obtains the actual cost and estimated cost of each estimated driving route, and then selects the estimated driving route with the minimum actual cost and estimated cost to generate an initial driving decision and execute it. When the user uploads operation instructions, the initial driving decision is automatically updated.