An intelligent interaction method and system for a car show, an electronic device, and a storage medium
By dividing the auto show exhibition area into independent interactive zones, collecting user action videos to identify operation intentions, and generating control commands to drive virtual vehicles, the problem of multiple people waiting in traditional auto shows is solved, and an efficient multi-user interactive experience is achieved.
Patent Information
- Application Number
- CN202510506814.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional auto shows often suffer from long waiting times due to limitations in user interaction devices, resulting in a reduced efficiency of the exhibition experience when multiple people are trying out the exhibits at the same time.
The auto show area is divided into multiple independent interactive zones. By collecting user action videos and identifying action trajectory features, control commands are generated to drive virtual vehicles, enabling personalized interaction for each user.
It breaks the limitations of traditional input devices, improves space utilization efficiency, ensures display quality when multiple users interact simultaneously, and enhances the overall display experience efficiency.
Smart Images

Figure CN120428857B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent auto show technology, specifically to an intelligent interactive method, system, electronic device, and storage medium for auto shows. Background Technology
[0002] With the rapid development of the automotive industry and the increasing demand from consumers for personalized vehicles, auto shows have become an important platform for automakers to showcase new models and promote new technologies. Traditional auto shows primarily rely on physical vehicle displays and sales staff explanations for interaction, allowing visitors to learn about the vehicles only through visual observation and limited tactile interaction. To enhance the display effect and user experience, auto shows have begun to introduce digital interactive technologies such as virtual reality in recent years, enabling users to more intuitively understand the appearance, functions, and performance characteristics of vehicles.
[0003] Currently, digital showrooms primarily use traditional input devices such as touchscreens, controllers, and wearable devices for interaction. For example, users can tap on a touchscreen to view different angles of a vehicle or use controller buttons to experience its features. Since each display terminal can only be operated by one user at a time, long waiting lines often form when there are large crowds at the exhibition, reducing the efficiency of the auto show's display experience. Summary of the Invention
[0004] This application provides an intelligent interactive method, system, electronic device, and storage medium for auto shows, which can improve the efficiency of the auto show display experience.
[0005] Firstly, this application provides an intelligent interaction method for auto shows, the method comprising:
[0006] The target auto show area was divided into multiple independent interactive zones;
[0007] When a user triggers an interaction request in the target interaction area, the target virtual vehicle is determined based on the interaction request, and each interaction area includes multiple virtual vehicles;
[0008] Collect video of the user's actions within the target interaction area, and determine the user's action trajectory features based on the video.
[0009] The motion trajectory features are matched with a preset motion feature library to identify the user's operation intention and generate corresponding control commands for the target virtual vehicle.
[0010] The control commands are transmitted to the control center of the target interactive area so that the control center drives the target virtual vehicle.
[0011] By adopting the technical scheme, the target car exhibition area is divided into multiple independent interaction areas, so that multiple users can operate in different interaction areas at the same time, avoiding mutual interference between users. When a user triggers an interaction request, a target virtual vehicle is determined based on the request. Since each interaction area includes multiple virtual vehicles, multiple users can experience in the same interaction area at the same time. Then, the action video of the user in the target interaction area is collected and the action trajectory feature is recognized. The operation intention of the user is recognized by matching the action trajectory feature with a preset action feature library, and the control instruction of the target virtual vehicle is generated. Finally, the control center of the target interaction area drives the target virtual vehicle to make an independent interaction response, realizing the exclusive interaction between each user and the corresponding virtual vehicle. This display method based on independent interaction areas not only breaks the limitation of traditional input devices and improves the space utilization efficiency of the exhibition area, but also ensures the display quality when multiple users interact at the same time through the independent interaction response mechanism, effectively improving the overall display experience efficiency of the car exhibition.
[0012] In a second aspect of the present application, an intelligent interaction system for a car exhibition is provided, which comprises:
[0013] An area division module is configured to divide a target car exhibition area into multiple independent interaction areas.
[0014] A vehicle determination module is configured to determine a target virtual vehicle based on an interaction request when a user triggers the interaction request in a target interaction area. Each interaction area includes multiple virtual vehicles.
[0015] A trajectory determination module is configured to collect the action video of a user in the target interaction area and determine the action trajectory feature of the user based on the action video.
[0016] An action matching module is configured to match the action trajectory feature with a preset action feature library, recognize the operation intention of the user, and generate the control instruction of the target virtual vehicle.
[0017] An interaction control module is configured to transmit the control instruction to the control center of the target interaction area, so that the control center drives the target virtual vehicle to realize an independent interaction response.
[0018] In a third aspect of the present application, a computer storage medium is provided, which stores a plurality of instructions. The instructions are suitable for being loaded and executed by a processor to perform the method steps described above.
[0019] In a fourth aspect of the present application, an electronic device is provided, which comprises a processor and a memory. The memory stores a computer program, which is suitable for being loaded and executed by the processor to perform the method steps described above.
[0020] To sum up, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0021] By dividing the target auto show area into multiple independent interaction areas, multiple users can simultaneously operate in different interaction areas, avoiding mutual interference between users. When a user triggers an interaction request, the target virtual vehicle is determined based on the request. Since each interaction area includes multiple virtual vehicles, multiple users can simultaneously experience in the same interaction area. Then, the action video of the user in the target interaction area is collected and the action trajectory feature is recognized. By matching with the preset action feature library, the operation intention of the user is recognized, and the control instruction of the target virtual vehicle is generated. Finally, the target virtual vehicle is driven by the control center of the target interaction area to make an independent interaction response, realizing the exclusive interaction between each user and the corresponding virtual vehicle. This display method based on independent interaction areas not only breaks the limitation of traditional input devices and improves the space utilization efficiency of the area, but also ensures the display quality when multiple users interact simultaneously through the independent interaction response mechanism, effectively improving the overall display experience efficiency of the auto show. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a flow diagram of an intelligent interaction method of an auto show provided by an embodiment of the present application;
[0023] Figure 2 is a module diagram of an intelligent interaction system of an auto show provided by an embodiment of the present application;
[0024] Figure 3 is a structural diagram of an electronic device provided by an embodiment of the present application.
[0025] BRIEF DESCRIPTION OF DRAWINGS DETAILED DESCRIPTION
[0026] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in conjunction with the drawings in the embodiments of the specification. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments.
[0027] In the description of the embodiments of the present application, the words "for example" or "for instance" are used to indicate an example, an instance, or an illustration. Any embodiment or design presented as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the words "for example" or "for instance" is intended to present related concepts in a specific manner.
[0028] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are used only for the purpose of description, and should not be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. The terms "include", "contain", "have" and their variants mean "include but are not limited to", unless otherwise specifically emphasized.
[0029] The intelligent interaction method of the car show related to the embodiments of the present application can be applied to large-scale display scenarios that need to meet the multi-user interaction requirements at the same time, and is especially suitable for application in digital display environments with limited space resources but high user experience requirements. For example, in the international car show, brand new car launch conference and other large-scale exhibition sites, there are usually multiple brand exhibition areas simultaneously exhibiting multiple new car models, and the number of on-site audience is large and the experience requirements are concentrated.
[0030] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments.
[0031] Please refer to Figure 1 , a flowchart of an intelligent interaction method of a car show is specifically proposed, which can be realized by relying on a computer program, can be realized by relying on a single-chip microcomputer, and can run on an intelligent interaction system of a car show. The computer program can be integrated in a computer device, or can run as an independent tool application. Specifically, the method includes steps 10 to 50, and the above steps are as follows:
[0032] Step 10: dividing a target car show area into a plurality of independent interaction areas.
[0033] The target car exhibition area in the embodiments of the present application refers to the overall space area for vehicle display, which is equipped with digital display equipment for presenting stereoscopic images of multiple virtual vehicles, and can be an independent exhibition hall, a specific area of an exhibition hall, or a temporary exhibition space. For example, it can be a brand exhibition area in a car exhibition site, or the entire display area of a digital experience center in a 4S store.
[0034] The interaction area in the embodiments of the present application refers to an independent operation space divided in the target car exhibition area, each interaction area is equipped with an image acquisition device and a virtual vehicle display device, and can support user interaction with virtual vehicles. The area has a clear spatial boundary, which can be divided by ground marking, projected boundary, or physical partition, to ensure that user operations in different interaction areas do not interfere with each other.
[0035] Specifically, first, the position information of each display car in the target car exhibition area is obtained, including the center point coordinates and the body orientation of the display car. Based on the position coordinates and the body size, the system calculates and determines the body projection area of each display car. To support multiple simultaneous interactions, a 4-5 meter preset distance is extended outward based on the body projection area to form an interaction buffer area, which can accommodate multiple people to experience at the same time. In each interaction area, the system simultaneously generates 3-5 independent virtual vehicle instances corresponding to the display car, and the initial positions of each virtual vehicle are relatively staggered and maintain a distance of more than 1 meter. When multiple users enter the same interaction area, the system automatically assigns each user a dedicated virtual vehicle instance, allowing multiple users to simultaneously interact with the same type of virtual vehicle without interference.
[0036] Based on the above embodiments, as an optional embodiment, the step of dividing the target car exhibition area into multiple independent interaction areas can further include the following steps:
[0037] Step 101: Obtain the position information of each display car in the target car exhibition area, including the center point coordinates and the body orientation of the display car.
[0038] Specifically, the three-dimensional images of the exhibition area are collected by multiple cameras and depth sensors arranged in the exhibition area, vehicle detection and contour recognition are performed on the collected three-dimensional images by computer vision algorithms, and the boundary point set of each display car is extracted. Based on the boundary point set, the center point coordinates of each display car are calculated using the minimum circumscribed rectangle algorithm. At the same time, by analyzing the direction of the connecting line of the front and rear bumpers of each display car, the body orientation angle is determined. The system records the center point coordinates (x, y) and the orientation angle θ of each display car in the spatial information database, and updates it in real time at a frequency of 10 Hz to ensure the accuracy of the position information. This way can accurately obtain the spatial pose of the display car.
[0039] Step 102: Determine the vehicle body projection area of each display vehicle based on the position information of each display vehicle and the vehicle body size.
[0040] Specifically, based on the obtained display vehicle position information, the system first reads the size parameters of the corresponding vehicle model in the vehicle three-dimensional model library, including vehicle length, vehicle width, wheelbase, etc. Combined with the center point coordinates and the orientation angle, the coordinates of the four corner points of the vehicle body are calculated using the affine transformation algorithm. By connecting these corner points, the rectangular projection contour of the vehicle body is obtained. Considering the operation requirements such as opening the door, the system expands 1.2 meters of operation space on both sides, and 0.8 meters of buffer distance in front and back, forming a complete vehicle body projection area. This area is marked by a white dotted line through the LED floor projection system, making the occupied space of the display vehicle clear and visible.
[0041] Step 103: Extend the interaction buffer area outward by a preset distance based on the vehicle body projection area, which is used to accommodate user interaction operations.
[0042] Specifically, to ensure that users have enough space to perform interaction operations, the system extends outward based on the vehicle body projection area using the equal distance expansion algorithm. The expansion distance is dynamically adjusted according to the interaction scene requirements: for the interaction scene of experiencing the vehicle's appearance, expand 2 meters outward; for the interaction scene that requires large amplitude actions, the expansion distance increases to 3 meters. The expansion algorithm uses a smooth transition with rounded corners to avoid sharp corners and improve user experience. The interaction buffer area is displayed with a blue gradient effect through the ground LED, and the display brightness is dynamically adjusted according to the number of users in the area. When the number of users approaches the upper limit, the color gradually changes to red as a prompt.
[0043] Step 104: Determine the combination area of the vehicle body projection area and the interaction buffer area as the interaction area of the corresponding display vehicle.
[0044] Specifically, the vehicle body projection area and the interaction buffer area are merged to generate the complete boundary of the interaction area. A spatial indexing algorithm is used to establish the management structure of the interaction area, supporting fast spatial query and collision detection. The system calculates the overlap between adjacent interaction areas in real time, and when overlap is detected, automatically adjusts the shape of the buffer area to eliminate interference. Each interaction area is equipped with independent computing resources and rendering units, responsible for handling all interaction requests and behavior control of virtual vehicles in that area. Through region division and resource isolation, stable and reliable multi-person simultaneous interaction experience is achieved.
[0045] Step 20: When there is a user triggered interaction request in the target interaction area, determine the target virtual vehicle based on the interaction request, and each interaction area includes multiple virtual vehicles.
[0046] Specifically, the user behavior is monitored in real time by the depth camera deployed in the interaction area. When the user is detected to enter the interaction area and stay for more than 5 seconds, the pre-judgment of the interaction request is triggered. The human posture recognition algorithm can be used to analyze the user's standing posture and gestures. When the user is recognized to face the display vehicle and make a preset trigger gesture (such as raising hands or pointing forward), it is confirmed that the trigger of the interaction request is successful. For this interaction request, the system first checks the number of existing virtual vehicle instances in the target interaction area, such as that each interaction area supports a maximum of 4 virtual vehicle instances to run simultaneously. These virtual vehicles are distributed in a diamond shape in space, and adjacent instances maintain a distance to ensure that the operations do not interfere with each other. The system maintains a virtual vehicle instance pool in real time, recording the usage state and associated user information of each instance.
[0047] When a new interaction request is received, the system preferentially selects the virtual vehicle closest to the user and in an idle state as the target virtual vehicle. If there is no idle instance nearby and the maximum instance number limit is not reached, a new virtual vehicle is created in real time as the target virtual vehicle. The newly created virtual vehicle gradually appears with a fade-in animation effect and automatically adjusts to a suitable observation angle. Once the target virtual vehicle is determined, the system immediately establishes a binding relationship between the user and the virtual vehicle, and displays an exclusive interaction boundary on the ground projection. This determination mechanism of the target virtual vehicle not only ensures the accuracy of the interaction, but also realizes efficient management of multi-user concurrent experience.
[0048] In another feasible embodiment, an intelligent tag recognition scheme can also be used to determine the target virtual vehicle. The user obtains an electronic name tag with a unique identifier when entering the exhibition area, and the RFID reader set in the interaction area determines the user's position by collecting the name tag signal. When the system detects the interaction request signal sent by the user, the determination process of the target virtual vehicle is started.
[0049] Step 30: Collect the action video of the user in the target interaction area, and determine the action trajectory features of the user based on the action video.
[0050] The action video refers to the image sequence continuously collected by the multiple high-speed cameras deployed in the target interaction area at a sampling frequency of 60 fps. These image sequences completely record the user's limb movements, position changes, and posture information, etc.
[0051] The action trajectory features refer to the set of user action feature parameters extracted from the action video.
[0052] Specifically, the high-speed camera of the target interaction area is used to collect the action video with a frame rate of 60 fps. The system first divides the target interaction area into multiple detection grids, and obtains the action process of the user in each detection grid through the cooperative work of multiple cameras. A deep learning model is used to extract 17 human body key points, including the three-dimensional space coordinates of the head, torso, limbs and other nodes, from the action video. Based on the position changes of these key points between adjacent frames, the system calculates the moving direction and speed, and then generates continuous action trajectory features. The features contain the spatial displacement and timing change information of the user's actions.
[0053] Based on the above embodiment, as another optional embodiment, the step of collecting the action video of the user in the target interaction area and determining the action trajectory features of the user based on the action video can further include the following steps:
[0054] Step 301: dividing the target interaction area into multiple detection grids.
[0055] Specifically, to achieve accurate action capture, the system uses an adaptive grid division algorithm to process the target interaction area. First, the boundary coordinates of the interaction area are obtained, and the entire area is divided into square detection grids with uniform size. The length of each grid is set to 0.5 meters. The system sets spatial reference markers at the four corners of each grid for subsequent coordinate conversion and position calibration. Adjacent grids share boundary marker points to form a continuous grid matrix. The system assigns a unique grid ID to each detection grid and establishes a grid index table for fast positioning and querying of action information in a specific grid.
[0056] Step 302: collecting the action video of the user in the target interaction area within a preset time period, and determining the original image sequence based on the action video, the original image sequence including the action process of the user in each detection grid.
[0057] Specifically, the high-speed camera in the target interaction area is used to synchronously collect the action video with a sampling frequency of 60 fps. The field of view of each camera is 120 degrees, and overlapping coverage is used to ensure no dead angle collection. The system sets a 3-second preset time period as the action collection window, during which 180 frames of images are continuously collected. The collected video data is processed by multiple cameras synchronously to remove redundant information and generate an original image sequence with a unified time reference. The system pre-processes the original image sequence, including denoising, illumination compensation and distortion correction, to ensure image quality. The processed image sequence clearly records the complete action process of the user in each detection grid.
[0058] Step 303: extracting the action node positions of the user from the original image sequence, the action node positions including the spatial coordinates of each human body key point in the detection grid.
[0059] Specifically, the original image sequence is analyzed by using a deep learning model to identify and extract 17 human body key points of the user, including the top of the head, the tip of the nose, the neck, the left and right shoulders, the left and right elbows, the left and right wrists, the center of the torso, the left and right hips, the left and right knees, and the left and right ankles. For each key point, the system calculates its accurate coordinates (x, y, z) in three-dimensional space. Through multi-camera perspective fusion and depth estimation algorithm, sub-centimeter positioning accuracy can be achieved. The spatial coordinates of these key points constitute the action node position information, which is used for subsequent trajectory analysis.
[0060] Step 304: Determine the moving direction and moving speed of the user's action based on the time sequence change of the spatial coordinates.
[0061] Specifically, based on the obtained action node position information, the system analyzes the change of the coordinates of each key point between adjacent frames. By calculating the coordinate difference vector, the instantaneous moving direction of each key point is obtained. At the same time, combined with the inter-frame time interval (1 / 60 seconds), the instantaneous moving speed of each key point is calculated. The system performs Kalman filtering on the speed data to eliminate sampling noise and obtain a smooth speed curve. By analyzing the trend of the speed curve, the system identifies acceleration, constant speed, and deceleration, etc. motion states, providing a basis for action feature extraction.
[0062] Step 305: Generate action trajectory features according to the moving direction and moving speed.
[0063] Specifically, in order to accurately generate action trajectory features that describe the complete action process of the user, the system first analyzes the change of the motion state at adjacent sampling times. The dot product of the moving direction vector of each two adjacent sampling times (t and t+1) is calculated to obtain the direction change angle θ(t). When θ(t) is greater than the preset threshold of 10 degrees, the time is marked as a potential turning point. The system uses a sliding window (window size of 5 frames) to smooth the angle sequence, eliminating jitter caused by sampling noise, and obtaining a stable direction change curve.
[0064] At the same time, the system analyzes the speed change between adjacent times. By calculating the speed difference ΔV(t), the motion state is judged: when ΔV(t)>0.2m / s, it is determined as an acceleration process, when ΔV(t)<-0.2m / s, it is determined as a deceleration process, and when |ΔV(t)|≤0.2m / s, it is determined as a constant speed process. The system applies a Kalman filter to the speed data to obtain a smooth speed trend curve.
[0065] Based on the distribution of the direction change angle, the system identifies the key turning points of the action. When the cumulative angle change exceeds 30 degrees, the point is marked as the boundary point of the action segment. In this way, the continuous action sequence can be decomposed into multiple basic action segments, each with a relatively stable motion direction.
[0066] For each action segment, the system extracts the following feature parameters: segment duration T, average speed V avg, maximum speed V max, speed standard deviation V std, start direction D start, end direction D end, cumulative turning angle θ total, average angular velocity ω avg. These parameters comprehensively describe the motion characteristics of the action segment.
[0067] Finally, the feature parameters of all action segments are organized in time sequence to construct a unified feature vector, that is, the action trajectory feature. The structure of the feature vector is: [segment number N, {T1, V avg1, V max1, V std1, D start1, D end1, θ total1, ω avg1},..., {TN, V avgN, V maxN, V stdN, D startN, D endN, θ totalN, ω avgN}]. This hierarchical feature representation method not only preserves the local action details, but also reflects the overall action structure.
[0068] Step 40: Match the action trajectory feature with the preset action feature library, identify the user's operation intention and generate the control instruction of the corresponding target virtual vehicle.
[0069] The preset action feature library in the embodiment of the application refers to a data set containing standardized action feature vectors.
[0070] The operation intention in the embodiment of the application refers to the specific interactive operation that the user expects to perform on the virtual vehicle identified by the system.
[0071] Specifically, the system pre-constructs an action feature library containing multiple groups of standard action feature templates, each template corresponding to a specific operation intention. When the action trajectory feature of the user is obtained, the dynamic time warping algorithm is used to calculate the similarity between it and each template in the action feature library. The system selects the operation intention corresponding to the template with the highest similarity and exceeding the preset threshold 0.85. Subsequently, the system queries the predefined intention-instruction mapping table to convert the identified operation intention into a specific control instruction of the target virtual vehicle, such as converting "forward" into "forward(speed=1)" instruction. This template matching based scheme can quickly and accurately identify the user's intention and generate the corresponding control instruction.
[0072] On the basis of the above embodiment, as another optional embodiment, the step of matching the action trajectory feature with the preset action feature library to identify the user's operation intention and generate the control instruction of the corresponding target virtual vehicle can further include the following steps:
[0073] Step 401: Extract key action nodes from action trajectory features.
[0074] Specifically, the extreme value analysis method is used to extract key action nodes from action trajectory features. First, the speed curve of each human key point is analyzed to identify local maximum and minimum points of speed, which usually correspond to the start, end or turning position of the action. The system sets the speed threshold to 0.5 m / s, and marks it as a potential key node when the speed change exceeds the threshold. At the same time, the acceleration curve is analyzed to identify the inflection points of acceleration positive and negative changes, which represent the transition of action state. By fusing the speed and acceleration feature points, the system finally determines the time stamp and spatial position of the key action node, providing basic data for subsequent feature vector construction.
[0075] Step 402: Construct action feature vector based on key action nodes, which contains spatial features and time sequence features of the action.
[0076] Specifically, based on the extracted key action nodes, the action feature vector containing spatial features and time sequence features is constructed. Spatial features include normalized coordinates of key points in three-dimensional space, relative distance and angle between adjacent key points. Time sequence features include time interval between key nodes, average speed and acceleration change rate. The system organizes these features into a fixed dimension vector, with the first 128 dimensions for spatial features and the last 64 dimensions for time sequence features. All feature values are normalized to ensure the value range is between [-1, 1]. This structured feature vector comprehensively describes the spatio-temporal characteristics of the action, facilitating subsequent pattern matching.
[0077] Step 403: Match the action feature vector with the action feature library to obtain the matching result.
[0078] Specifically, a multi-level matching strategy is used to match the action feature vector with the action feature library. First, the Local Sensitive Hashing (LSH) algorithm is used to quickly filter out candidate templates in the feature library, reducing the search space. Then, the Dynamic Time Warping (DTW) algorithm is applied to the candidate templates to calculate the similarity score between the feature vector and each template. The system sets the similarity threshold to 0.85, and selects the matching results with scores higher than the threshold. For multiple matching results, the system selects the optimal match based on a weighted voting mechanism, considering the similarity score and historical matching frequency. This gradual matching scheme ensures the accuracy of the match and improves the processing efficiency.
[0079] Step 404: Determine the user's operation intention based on the matching result, including operation direction and operation amplitude.
[0080] Specifically, the pre-defined operation type label is extracted from the best matched template, and then the direction component in the feature vector is analyzed to calculate the unit vector of the main motion direction to determine the operation direction. Meanwhile, the operation amplitude is determined by analyzing the key point displacement, such as the rotation angle is determined by calculating the included angle of the start and end points for the rotation operation, and the moving distance is determined by calculating the modulus of the displacement vector for the moving operation. The system integrates these parameters into a complete operation intention description, including operation type, direction and amplitude information.
[0081] Step 405: generating control instructions of the corresponding target virtual vehicle according to the operation intention.
[0082] Specifically, the system pre-establishes a mapping table of operation intention and control instruction, and generates control instructions in a standard format based on the recognized operation intention. The operation intention is parsed into basic operation units, such as rotation axis, rotation direction and angle parameters for the rotation operation. Then, these parameters are converted into JSON format control instructions according to the mapping rules, including instruction type, target object ID, action parameters and other fields. For example, the operation of rotating 90 degrees clockwise will generate the instruction: {"type":"rotate","target":"car_001","axis":"y","angle":90,"duration":2000}.
[0083] On the basis of the above embodiment, as another optional embodiment, the step of determining the operation intention of the user based on the matching result can further include the following steps:
[0084] Step 4041: when the matching result is a matching success, the operation parameters of the pre-set action feature corresponding to the matching result are taken as the operation intention of the user.
[0085] Specifically, when the similarity between the action feature vector and a certain pre-set action feature in the feature library exceeds 0.85, the system determines that the matching is successful. At this time, the system directly extracts the operation parameters from the matched pre-set action feature record, including operation type (such as "rotate", "move", "zoom"), operation direction (such as rotation axis vector, moving direction vector) and operation amplitude (such as rotation angle, moving distance, zoom ratio).
[0086] Step 4042: when the matching result is a matching failure, multiple candidate actions with similarity greater than a pre-set similarity to the action feature vector are selected from the action feature library.
[0087] Specifically, when the highest matching similarity is lower than 0.85, the system starts the candidate action screening mechanism. The system uses the cosine similarity algorithm to calculate the similarity of the action feature vector and all preset actions in the feature library, and selects the actions with a similarity greater than 0.7 as the candidate set. To improve retrieval efficiency, the system first uses local sensitive hashing (LSH) pre-screening to divide the feature space into blocks and only performs accurate matching in the relevant blocks. The system also considers the spatiotemporal feature similarity and context relevance of the action to ensure that the candidate action is consistent with the current interaction scenario.
[0088] Step 4043: Display each candidate action in the form of an interaction menu in the target interaction area, and arrange the candidate actions in the interaction menu from high to low similarity.
[0089] Specifically, a semi-transparent arc-shaped interaction menu is generated in the target interaction area, and the menu is located 0.8 meters in front of the user's line of sight. The candidate actions are displayed in the form of icons and text, and each option includes an action type icon, a brief description text, and a similarity percentage. The system arranges the candidate options in descending order of similarity, with the most similar option in the center of the menu and the other options expanding in a fan shape. The menu uses a gradient transparency effect to ensure that it does not block the user's view of the virtual vehicle. Each option is equipped with visual feedback effects and appears in a highlighted state when the user's line of sight or gesture hovers over it.
[0090] Step 4044: When a selection operation of the user on the interaction menu is detected within a preset time period, determine the candidate action corresponding to the selection operation as the user's operation intent.
[0091] Specifically, the system monitors the user's operation on the interaction menu within a preset time period of 5 seconds, supporting two selection methods: gesture clicking and gaze confirmation. Gesture clicking is triggered by detecting the user's hand staying and pressing action in the option area; gaze confirmation requires the user's line of sight to stay on the option for more than 0.8 seconds. When a valid selection operation is detected, the system immediately hides the menu and extracts the operation parameters of the selected candidate action as the operation intent. The system also records this selection to the user feedback database for optimizing future matching strategies.
[0092] Step 4045: When no selection operation of the user on the interaction menu is detected within a preset time period, enter the intelligent assistant mode, which is used to provide interaction operation guidance and action demonstration.
[0093] Specifically, when the selection operation is not detected for more than 5 seconds, the system smoothly transitions to the intelligent assistant mode. First, a personified virtual assistant image is displayed in the interaction area, and the assistant uses natural language to introduce the currently available interaction operations. The system also displays action demonstration animations in the virtual space, showing standard action postures through a semi-transparent human outline. The assistant will observe the user's actions in real time and provide immediate feedback when the user tries to imitate, pointing out the key points of the action and suggestions for improvement. This interactive guidance helps users quickly master the correct operation method. The intelligent assistant mode continues until the user successfully completes an effective operation or actively exits.
[0094] Step 50: transmit the control instruction to the control center of the target interaction area to drive the target virtual vehicle.
[0095] The control center in the embodiments of the present application refers to a local processing unit independently configured for each interaction area, responsible for managing the interaction process within the interaction area.
[0096] Specifically, the system allocates an independent communication channel for each interaction area and uses the UDP protocol to transmit JSON format control instructions to the control center of the corresponding target interaction area. After receiving the instruction, the control center first checks the running state and resource occupation of the virtual vehicle in the area to ensure that the new instruction can be executed. Subsequently, the control center parses the control instruction into transformation parameters supported by the rendering engine of the area, and executes the form adjustment of the target virtual vehicle through the graphics processing unit in the area. This region-based distributed control method ensures independent operation between different interaction areas and avoids interaction interference.
[0097] On the basis of the above-mentioned embodiments, as an optional embodiment, an intelligent interaction method for a car show can further include the following processes:
[0098] Specifically, the number of users is detected in real time by the depth camera deployed in each interaction area. For example, a human detection algorithm based on YOLOv5 can be used to perform picture analysis every 100 ms, extracting the position and ID features of each detected human target. The system uses the DeepSORT algorithm for target tracking, recording the entry time and exit time of each user. Within a preset statistical period of 60 seconds, the system accumulatively calculates the residence time of each user, and adds the residence time of all users in the same interaction area to obtain the total residence time of the area. For example, the system sets the user number threshold to 3 and the total residence time threshold to 5 minutes. When the real-time number of users in a certain interaction area exceeds 3 or the total residence time exceeds 5 minutes, the area is marked as a hot interaction area, triggering the area expansion mechanism.
[0099] When the hotspot interaction area is identified, the system performs area expansion, each interaction area has reserved 1.5 times expandable space when planning, and the reserved spaces maintain a 1-meter safety distance between adjacent interaction areas, ensuring that expansion does not interfere with each other. The system expands in two stages according to the number of users: when the number of users is 3, the area is expanded by 1.2 times; when the number of users is 4, the area is expanded by 1.5 times. When the number of users in the area exceeds 4, the system suggests additional users to go to other idle interaction areas through voice prompts and display screen prompts.
[0100] The expanded area boundary is identified by a blue laser line, and the line width is kept at 2 cm. The system adjusts the display brightness of the boundary line according to the number of users: the basic brightness is set to 50 nit, which is increased to 75 nit when there are 3 people, and increased to 100 nit when there are 4 people. The boundary line uses a breathing animation effect, and the brightness periodically changes between 80%-100% of the set value with a period of 2 seconds. At the same time, the system synchronously adjusts the interaction devices in the expanded area: recalibrates the camera coverage, adjusts the projection parameters of the display device, and ensures the quality of the interaction experience in the expanded space. When the number of users drops to 2 or less and lasts for 3 minutes, the system automatically restores the area to its original size. This expansion mechanism strictly limits the maximum number of users, ensuring the quality of the interaction experience and maintaining the orderly operation of the exhibition area.
[0101] See Figure 2 , a module schematic diagram of an intelligent interaction system of a car exhibition provided for an embodiment of the present application, wherein the system comprises:
[0102] An area division module, configured to divide a target car exhibition area into a plurality of independent interaction areas;
[0103] A vehicle determination module, configured to determine a target virtual vehicle based on an interaction request when a user triggers the interaction request in a target interaction area, and each of the interaction areas includes a plurality of virtual vehicles;
[0104] A trajectory determination module, configured to collect action video of a user in the target interaction area, and determine action trajectory features of the user based on the action video;
[0105] An action matching module, configured to match the action trajectory features with a preset action feature library, identify the operation intention of the user, and generate a control instruction for the target virtual vehicle;
[0106] An interaction control module, configured to transmit the control instruction to a control center of the target interaction area, so that the control center drives the target virtual vehicle.
[0107] Optionally, the region division module is further configured to acquire position information of each display vehicle in the target vehicle exhibition area, the position information comprising a center point coordinate and a vehicle body orientation of the display vehicle;
[0108] based on the position coordinate and the vehicle body size of each display vehicle, determine a vehicle body projection region of each display vehicle;
[0109] based on the vehicle body projection region, expand outward by a preset distance to obtain an interactive buffer region, the interactive buffer region being configured to accommodate user interaction;
[0110] determine a combination region of the vehicle body projection region and the interactive buffer region as an interactive region of the corresponding display vehicle.
[0111] Optionally, the trajectory determination module is further configured to divide the target interactive region into a plurality of detection grids;
[0112] acquire motion video of a user in the target interactive region within a preset time period, and determine an original image sequence based on the motion video, the original image sequence comprising a motion process of the user in each detection grid;
[0113] extract motion node positions of the user from the original image sequence, the motion node positions comprising spatial coordinates of each human body key point in the detection grid;
[0114] determine a moving direction and a moving speed of the user motion based on a time sequence change of the spatial coordinates;
[0115] generate the motion trajectory feature according to the moving direction and the moving speed.
[0116] Optionally, the trajectory determination module is further configured to calculate a direction change angle of the user motion based on a moving direction of adjacent time instants, the direction change angle representing a turning feature of the motion;
[0117] determine a speed change trend of the user motion based on a moving speed of adjacent time instants, the speed change trend comprising an acceleration process and a deceleration process;
[0118] divide the user motion into a plurality of motion segments according to a distribution of the direction change angle;
[0119] extract a motion feature parameter of each motion segment, and generate the motion trajectory feature based on the motion feature parameter of each motion segment.
[0120] Optionally, the motion matching module is further configured to extract a key motion node from the motion trajectory feature;
[0121] construct an action feature vector based on the key action node, the action feature vector comprising spatial features and timing features of the action;
[0122] match the action feature vector with the action feature library to obtain a matching result;
[0123] determine an operation intention of the user based on the matching result, the operation intention comprising an operation direction and an operation amplitude;
[0124] generate a control instruction of the target virtual vehicle corresponding to the operation intention.
[0125] Optionally, the action matching module is further configured to, when the matching result is a matching success, take an operation parameter of a preset action feature corresponding to the matching result as the operation intention of the user.
[0126] when the matching result is a matching failure, select a plurality of candidate actions with a similarity greater than a preset similarity from the action feature library;
[0127] display each of the candidate actions in the form of an interactive menu in the target interactive area, the candidate actions in the interactive menu being arranged from high to low according to the similarity;
[0128] when a selection operation of the user on the interactive menu is detected within a preset time period, determine the candidate action corresponding to the selection operation as the operation intention of the user;
[0129] when a selection operation of the user on the interactive menu is not detected within a preset time period, enter an intelligent assistant mode, the intelligent assistant mode being configured to provide interactive operation guidance and action demonstration.
[0130] Optionally, the intelligent interactive system of the car show further comprises a region adjustment module configured to detect a number of users in each of the interactive areas in real time, and to count a total residence time of all users in each of the interactive areas within a preset period.
[0131] when there is a hotspot interactive area with a number of users greater than a number threshold and / or a total residence time greater than a time threshold, expand the hotspot interactive area according to the number of users and a preset expansion strategy to obtain an expanded interactive area, a boundary of the expanded interactive area being displayed in a preset color, and a display brightness of the boundary being enhanced with an increase in the number of users.
[0132] It should be noted that the system provided in the above embodiments is only used as an example to divide the above functional modules when realizing the functions thereof, and in actual application, the above functions can be completed by different functional modules according to the needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be described here.
[0133] The computer storage medium provided in the embodiments of the present application can store a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above-mentioned method for intelligent interaction of a car show. The specific implementation process can be referred to the specific description of the above-mentioned method embodiments, which will not be described here.
[0134] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 is a structural schematic diagram of an electronic device disclosed in the embodiments of the present application. The electronic device 300 can include at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0135] The communication bus 302 is used to realize the connection and communication between the components.
[0136] The user interface 303 can include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 can also include a standard wired interface and a wireless interface.
[0137] The network interface 304 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0138] The processor 301 can include one or more processing cores. The processor 301 connects various parts within the server through various interfaces and lines, performs various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and calling data stored in the memory 305. Alternatively, the processor 301 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 301 can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 301, but can be realized by a separate chip.
[0139] The memory 305 can include a random access memory (RAM) and a read-only memory (ROM). Alternatively, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 305 can also be at least one storage device located away from the aforementioned processor 301. Referring to Figure 3 The memory 305 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an application program of the intelligent interaction method of the car show.
[0140] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input, and obtain data input by the user; and the processor 301 can be used to invoke an application program stored in the memory 305 and storing a location guiding method for blind people to travel, which, when executed by one or more processors 301, causes the electronic device 300 to perform the method of one or more of the above-described embodiments. It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0141] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0142] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner for actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical or other forms.
[0143] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0144] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0145] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: a U disk, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0146] The above are only exemplary embodiments of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practicing the true principles of the present disclosure.
[0147] The present application is intended to cover any variations, uses or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field not described in the present disclosure. The specification and examples are only considered as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. An intelligent interaction method for a car show, characterized in that, The method comprises: dividing a target auto show exhibition area into multiple independent interaction areas, the target auto show exhibition area being a whole space area for vehicle display, the area being equipped with digital display equipment for presenting stereoscopic images of multiple virtual vehicles; in each interaction area, the system simultaneously generates multiple independent virtual vehicle instances corresponding to the display vehicle, and when multiple users enter the same interaction area, the system automatically allocates a dedicated virtual vehicle instance to each user, so that the multiple users can simultaneously have independent interaction experience on the same type of virtual vehicle without interference; when there is a user-triggered interaction request in a target interaction area, determining a target virtual vehicle based on the interaction request, each interaction area including multiple virtual vehicles; collecting action video of the user in the target interaction area, and determining action trajectory features of the user based on the action video; matching the action trajectory features with a preset action feature library, identifying the operation intention of the user, and generating control instructions for the target virtual vehicle; transmitting the control instructions to a control center of the target interaction area, so that the control center drives the target virtual vehicle; The method further comprises: real-time detection of the number of users in each interaction area, and statistics of the total residence time of all users in each interaction area within a preset period; when there is a hotspot interaction area with a user number greater than a number threshold and / or a total residence time greater than a time threshold, expanding the hotspot interaction area according to the user number and a preset expansion strategy to obtain an expanded interaction area, the boundary of the expanded interaction area being displayed in a preset color, and the display brightness of the boundary being enhanced with the increase of the user number.
2. The intelligent interaction method for car show according to claim 1, characterized in that, The method further comprises: obtaining position information of each display vehicle in the target auto show exhibition area, the position information including the center point coordinates and the body orientation of the display vehicle; determining the body projection area of each display vehicle based on the position information and the body size of each display vehicle; extending a preset distance outward based on the body projection area to obtain an interaction buffer area, the interaction buffer area being used to accommodate user interaction operations; determining the combination area of the body projection area and the interaction buffer area as the interaction area of the corresponding display vehicle.
3. The intelligent interaction method of car show according to claim 1, characterized in that, The method further comprises: dividing the target interaction area into multiple detection grids; collecting action video of the user in the target interaction area within a preset time period, determining an original image sequence based on the action video, the original image sequence including the action process of the user in each detection grid; extracting the action node position of the user from the original image sequence, the action node position including the spatial coordinates of each human body key point in the detection grid; determining the moving direction and moving speed of the user's action based on the time sequence change of the spatial coordinates; generating the action trajectory features according to the moving direction and the moving speed.
4. The intelligent interaction method of car show according to claim 3, characterized in that, The generating the action trajectory feature according to the moving direction and the moving speed comprises: calculating a direction change angle of the user action based on the moving direction of the adjacent time, the direction change angle representing a turning feature of the action; determining a speed change trend of the user action based on the moving speed of the adjacent time, the speed change trend comprising an acceleration process and a deceleration process; dividing the user action into a plurality of action segments according to the distribution of the direction change angle and the speed change trend; extracting an action feature parameter of each of the action segments, and generating the action trajectory feature based on the action feature parameter of each of the action segments.
5. The intelligent interaction method of car show according to claim 1, characterized in that, The matching the action trajectory feature with the preset action feature library, identifying the operation intention of the user and generating the control instruction of the target virtual vehicle corresponding to the target virtual vehicle comprises: extracting a key action node from the action trajectory feature; constructing an action feature vector based on the key action node, the action feature vector containing spatial features and time sequence features of the action; matching the action feature vector with the action feature library to obtain a matching result; determining the operation intention of the user based on the matching result, the operation intention comprising an operation direction and an operation amplitude; generating the control instruction of the target virtual vehicle corresponding to the target virtual vehicle according to the operation intention.
6. The intelligent interaction method of car show according to claim 5, characterized in that, The determining the operation intention of the user based on the matching result comprises: when the matching result is a matching success, taking an operation parameter of a preset action feature corresponding to the matching result as the operation intention of the user; when the matching result is a matching failure, selecting a plurality of candidate actions with a similarity greater than a preset similarity to the action feature vector from the action feature library; displaying each of the candidate actions in the form of an interactive menu in the target interactive area, the candidate actions in the interactive menu being arranged from high to low in similarity; when a selection operation of the user on the interactive menu is detected within a preset time period, determining the candidate action corresponding to the selection operation as the operation intention of the user; when no selection operation of the user on the interactive menu is detected within a preset time period, entering an intelligent assistant mode, the intelligent assistant mode being used to provide interactive operation guidance and action demonstration.
7. An intelligent interaction system for a car show, characterized in that, An intelligent interaction method for performing a car show as claimed in claim 1, the system comprising: a region division module configured to divide a target car show area into a plurality of independent interactive areas; a vehicle determination module configured to determine a target virtual vehicle based on an interaction request when the interaction request is triggered by a user in a target interactive area, each of the interactive areas comprising a plurality of virtual vehicles; a trajectory determination module configured to collect an action video of the user in the target interactive area, and determine an action trajectory feature of the user based on the action video; an action matching module configured to match the action trajectory feature with a preset action feature library, identify an operation intention of the user, and generate a control instruction of the target virtual vehicle corresponding to the target virtual vehicle; an interaction control module configured to transmit the control instruction to a control center of the target interactive area, so that the control center drives the target virtual vehicle.
8. A computer-readable storage medium, characterized in that, A computer readable storage medium stores a plurality of instructions adapted to be loaded and executed by a processor to perform the method of any of claims 1-6.
9. An electronic device, comprising: An electronic device includes a processor, a memory, a user interface, and a network interface, the memory is configured to store instructions, the user interface and the network interface are configured to communicate with other devices, and the processor is configured to execute the instructions stored in the memory to cause the electronic device to perform the method of any of claims 1-6.
Citation Information
Patent Citations
Vehicle information display method and device, storage medium and program product
CN119359401A
Intelligent AI holographic projection method and system based on artificial intelligence
CN119723103A