Intelligent interaction method and system for vehicle exhibition, electronic equipment and storage medium
By dividing independent interaction areas in the auto show exhibition area, collecting user action videos to identify operation intentions and generating virtual vehicle control instructions, the problem of restricted user interaction in traditional auto show is solved, and the display experience of efficient interaction between multiple users is achieved.
Patent Information
- Application Number
- CN202510506814.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In traditional auto show, user interaction methods are limited by physical equipment, resulting in a long wait team formed under the condition of large traffic, reducing the efficiency of the display experience.
The auto show area is divided into multiple independent interactive areas, and by collecting user action videos to identify operation intentions and generating control instructions for virtual vehicles, each user can realize exclusive interaction with their corresponding virtual vehicles.
It breaks the limitations of traditional input devices, improves the space utilization efficiency of the exhibition area and the display quality of multiple users' simultaneous interaction, and improves the overall display experience efficiency of the auto show.
Smart Images

Figure CN120428857A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of smart auto shows, and in particular to an intelligent interaction method, system, electronic device, and storage medium for auto shows. Background Art
[0002] With the rapid development of the automotive industry and increasing consumer demand for personalized vehicles, auto shows have become a crucial platform for automakers to showcase new models and promote new technologies. Traditional auto shows primarily rely on physical vehicle displays and sales staff explanations, limiting visitors' understanding of the vehicles to visual observation and limited touch. To enhance presentation and user experience, auto shows have recently begun to incorporate digital interactive technologies such as virtual reality, allowing users to more intuitively understand the vehicle's appearance, functions, and performance characteristics.
[0003] Currently, digital showrooms primarily utilize traditional input devices such as touchscreens, controllers, and wearable devices for interaction. For example, users can tap on a touchscreen to view a vehicle from different angles or press buttons on a controller to experience its features. Because each display terminal can only be operated by one user at a time, long waits often form during high-traffic exhibitions, reducing the efficiency of the auto show experience. Summary of the Invention
[0004] The present application provides an intelligent interaction method, system, electronic device and storage medium for an auto show, which can improve the display experience efficiency of the auto show.
[0005] In a first aspect, the present application provides an intelligent interaction method for an auto show, the method comprising: Divide the target auto show area into multiple independent interactive areas; When a user triggers an interaction request in a target interaction area, determining a target virtual vehicle based on the interaction request, wherein each interaction area includes a plurality of virtual vehicles; Collecting a motion video of a user in the target interaction area, and determining a motion trajectory feature of the user based on the motion video; Matching the motion trajectory features with a preset motion feature library, identifying the user's operation intention and generating corresponding control instructions for the target virtual vehicle; The control instruction is transmitted to a control center of the target interaction area, so that the control center drives the target virtual vehicle.
[0006] By adopting the above technical solution, the target auto show area is divided into multiple independent interactive zones, allowing multiple users to operate simultaneously in different interactive zones, thus avoiding mutual interference between users. When a user triggers an interaction request, the target virtual vehicle is determined based on the request. Since each interactive zone includes multiple virtual vehicles, multiple users can experience the same interactive zone simultaneously. Then, the user's action video in the target interactive zone is collected and the motion trajectory features are identified. The user's operation intention is identified by matching it with a preset motion feature library, and then control instructions for the target virtual vehicle are generated. Finally, the control center of the target interactive zone drives the target virtual vehicle to make an independent interactive response, realizing exclusive interaction between each user and their corresponding virtual vehicle. This display method based on independent interactive zones not only breaks the limitations of traditional input devices and improves the space utilization efficiency of the exhibition area, but also ensures the display quality when multiple users interact simultaneously through an independent interactive response mechanism, effectively improving the overall display experience efficiency of the auto show.
[0007] In a second aspect of the present application, an intelligent interactive system for an auto show is provided, the system comprising: The area division module is used to divide the target auto show area into multiple independent interactive areas; a vehicle determination module, configured to, when a user triggers an interaction request in a target interaction area, determine a target virtual vehicle based on the interaction request, wherein each interaction area includes a plurality of virtual vehicles; a trajectory determination module, configured to collect a motion video of a user in the target interaction area and determine motion trajectory features of the user based on the motion video; An action matching module is used to match the action trajectory features with a preset action feature library, identify the user's operation intention and generate corresponding control instructions for the target virtual vehicle; The interactive control module is used to transmit the control instruction to the control center of the target interactive area, so that the control center drives the target virtual vehicle to achieve an independent interactive response.
[0008] In a third aspect of the present application, a computer storage medium is provided. The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above method steps.
[0009] In a fourth aspect of the present application, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
[0010] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: This application divides the target auto show area into multiple independent interactive areas, allowing multiple users to operate in different interactive areas at the same time, thus avoiding mutual interference between users. When a user triggers an interactive request, the target virtual vehicle is determined based on the request. Since each interactive area includes multiple virtual vehicles, multiple users can experience it simultaneously in the same interactive area. Then, the action video of the user in the target interactive area is collected and the action trajectory features are identified. The user's operation intention is identified by matching with the preset action feature library, and then the control instructions of the target virtual vehicle are generated. Finally, the target virtual vehicle is driven by the control center of the target interactive area to make an independent interactive response, thus realizing exclusive interaction between each user and its corresponding virtual vehicle. This display method based on independent interactive areas not only breaks the limitations of traditional input devices and improves the space utilization efficiency of the exhibition area, but also ensures the display quality when multiple users interact at the same time through an independent interactive response mechanism, effectively improving the overall display experience efficiency of the auto show. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a flow chart of an intelligent interaction method for an auto show provided in an embodiment of the present application; Figure 2 This is a module diagram of an intelligent interactive system for an auto show provided by an embodiment of the present application; Figure 3 This is a structural diagram of an electronic device provided in an embodiment of the present application.
[0012] Description of reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION
[0013] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0014] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0015] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0016] The intelligent interaction method for auto shows involved in the embodiments of this application can be applied to large-scale exhibition scenarios that need to meet the interactive needs of multiple users simultaneously. It is particularly suitable for use in digital exhibition environments with limited space resources but high user experience requirements. For example, at large exhibitions such as international auto shows and brand new car launches, multiple brand exhibition areas often display multiple new models simultaneously, attracting large audiences with concentrated experience requirements.
[0017] The following will provide a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0018] Please refer to Figure 1 , a flowchart of an intelligent interactive method for an auto show is proposed. The method can be implemented by a computer program, a single-chip microcomputer, or run on an intelligent interactive system for an auto show. The computer program can be integrated into a computer device or run as an independent tool application. Specifically, the method includes steps 10 to 50, which are as follows: Step 10: Divide the target auto show area into multiple independent interactive areas.
[0019] In this embodiment of the present application, the target auto show exhibition area refers to the entire spatial area used for vehicle display, equipped with digital display equipment for presenting 3D images of multiple virtual vehicles. This area can be an independent exhibition hall, a specific area of a pavilion, or a temporary display space. For example, at an auto show, this could be the booth area of a specific brand, or at a 4S dealership, it could be the entire display area of a digital experience center.
[0020] In this embodiment, interactive zones refer to independent operating spaces within the target auto show area. Each interactive zone is equipped with an image acquisition device and a virtual vehicle display device, enabling users to interact with the virtual vehicle. These zones have clear spatial boundaries and can be demarcated using floor markings, projected boundaries, or physical partitions to ensure that users operating within different interactive zones do not interfere with each other.
[0021] Specifically, the system first obtains the position information of each display car in the target auto show area, including the coordinates of the center point and the orientation of the body of the display car. Based on the position coordinates and the body size, the system calculates and determines the body projection area of each display car. In order to support simultaneous interaction among multiple people, an interactive buffer area is formed by expanding outward by a preset distance of 4-5 meters based on the body projection area. This area can accommodate multiple people to experience at the same time. In each interactive area, the system simultaneously generates 3-5 independent virtual vehicle instances corresponding to the display car. The initial position of each virtual vehicle is relatively staggered and maintains a distance of more than 1 meter. When multiple users enter the same interactive area, the system automatically assigns an exclusive virtual vehicle instance to each user, allowing multiple users to simultaneously and independently interact with the same model of virtual vehicle without interfering with each other.
[0022] Based on the above embodiment, as an optional embodiment, the step of dividing the target auto show area into multiple independent interactive areas may further include the following steps: Step 101: Obtaining the position information of each display vehicle in the target auto show area, where the position information includes the center point coordinates and the body orientation of the display vehicle.
[0023] Specifically, multiple cameras and depth sensors installed within the exhibition area capture a three-dimensional image of the exhibition area. Computer vision algorithms then perform vehicle detection and contour recognition on these captured 3D images, extracting a set of boundary points for each display vehicle. Based on this set of boundary points, the minimum enclosing rectangle algorithm is used to calculate the coordinates of each display vehicle's center point. Furthermore, the orientation angle of each display vehicle is determined by analyzing the direction of the line connecting the front and rear bumpers. The system records the center point coordinates (x, y) and orientation angle θ of each display vehicle in a spatial information database and updates it in real time at a 10Hz frequency to ensure the accuracy of the positional information. This method accurately captures the spatial posture of each display vehicle.
[0024] Step 102: Determine the body projection area of each display vehicle based on the position information and body size of each display vehicle.
[0025] Specifically, based on the acquired display vehicle's location information, the system first retrieves the corresponding model's dimensional parameters from the 3D vehicle model library, including length, width, and wheelbase. Combining the center point coordinates and orientation angle, an affine transformation algorithm is used to calculate the coordinates of the vehicle's four corner points. By connecting these corner points, a rectangular projection outline of the vehicle body is obtained. To accommodate operational needs such as door opening, the system extends the operating space by 1.2 meters on each side and a 0.8-meter buffer distance at the front and rear to create a complete vehicle projection area. This area is marked with a white dashed line using an LED floor projection system, clearly indicating the space occupied by the display vehicle.
[0026] Step 103: Taking the vehicle body projection area as a reference, expand outward by a preset distance to obtain an interactive buffer area, where the interactive buffer area is used to accommodate user interactive operations.
[0027] Specifically, to ensure users have sufficient space for interactive operations, the system uses an equidistant expansion algorithm to extend the projection area of the vehicle body. This expansion distance is dynamically adjusted based on the needs of the interactive scenario: for interactive scenarios that require experiencing the vehicle's exterior, the distance is extended outward by 2 meters; for interactive scenarios requiring larger movements, the distance is increased to 3 meters. The expansion algorithm uses a smooth transition and rounded corners to avoid sharp corners and enhance the user experience. The interactive buffer area is displayed using a blue gradient effect via ground-based LEDs, and the display brightness is dynamically adjusted based on the number of users in the area. When the number of users approaches the upper limit, the color gradually changes to red as a prompt.
[0028] Step 104: Determine the combined area of the vehicle body projection area and the interaction buffer area as the interaction area corresponding to the display vehicle.
[0029] Specifically, the vehicle projection area and the interaction buffer area are merged to generate a complete interaction area boundary. A spatial indexing algorithm is used to establish a management structure for interaction areas, supporting fast spatial queries and collision detection. The system calculates overlap between adjacent interaction areas in real time and, when overlap is detected, automatically adjusts the shape of the buffer area to eliminate interference. Each interaction area is equipped with independent computing resources and rendering units, responsible for processing all interaction requests and controlling the behavior of the virtual vehicle within that area. This regional division and resource isolation ensures a stable and reliable multi-person simultaneous interaction experience.
[0030] Step 20: When a user triggers an interaction request in a target interaction area, a target virtual vehicle is determined based on the interaction request. Each interaction area includes multiple virtual vehicles.
[0031] Specifically, depth cameras deployed in the interaction area monitor user behavior in real time. When a user enters the interaction area and remains there for more than five seconds, a pre-judgment of the interaction request is triggered. A human posture recognition algorithm can be used to analyze the user's posture and gestures. When the user is identified as facing the display vehicle and making a preset trigger gesture (such as raising a hand or pointing forward), the interaction request is confirmed to have been successfully triggered. In response to this interaction request, the system first checks the number of existing virtual vehicle instances within the target interaction area. For example, each interaction area supports up to four virtual vehicle instances running simultaneously. These virtual vehicles are spatially distributed in a diamond shape, with adjacent instances maintaining spacing to ensure that operations do not interfere with each other. The system maintains a pool of virtual vehicle instances in real time, recording the usage status and associated user information of each instance.
[0032] When a new interaction request is received, the system prioritizes the idle virtual vehicle closest to the user as the target virtual vehicle. If there are no idle instances nearby and the maximum number of instances has not been reached, a new virtual vehicle is created in real time as the target virtual vehicle. The newly created virtual vehicle gradually appears using a fade-in animation and automatically adjusts to the appropriate viewing angle. Once the target virtual vehicle is determined, the system immediately establishes a binding relationship between the user and the virtual vehicle and projects a dedicated interaction boundary on the ground. This mechanism for determining the target virtual vehicle ensures both accurate interaction and efficient management of concurrent multi-user experiences.
[0033] In another feasible embodiment, a smart tag recognition solution can be used to determine the target virtual vehicle. Upon entering the exhibition area, users are given a unique electronic badge. An RFID reader located within the interactive area detects the badge signal and determines the user's location. When the system detects an interaction request signal from the user, the target virtual vehicle determination process is initiated.
[0034] Step 30: Collect the user's action video in the target interaction area, and determine the user's action trajectory characteristics based on the action video.
[0035] Action videos are image sequences continuously captured at 60fps by multiple high-speed cameras deployed in the target interaction area. These image sequences fully record the user's body movements, position changes, and posture information.
[0036] Action trajectory features refer to the set of user action feature parameters extracted from action videos.
[0037] Specifically, a high-speed camera in the target interaction area captures action videos at a frame rate of 60fps. The system first divides the target interaction area into multiple detection grids. Using the collaborative work of multiple cameras, the system captures the user's movements within each detection grid. A deep learning model is then used to extract 17 key human points from the action videos, including the 3D spatial coordinates of nodes such as the head, torso, and limbs. Based on the positional changes of these key points between adjacent frames, the system calculates the movement direction and speed, generating a continuous motion trajectory feature. This feature captures the spatial displacement and temporal changes of the user's movements.
[0038] Based on the above embodiment, as another optional embodiment, the step of collecting the user's action video in the target interaction area and determining the user's action trajectory characteristics based on the action video may further include the following steps: Step 301: Divide the target interaction area into multiple detection grids.
[0039] Specifically, to achieve accurate motion capture, the system uses an adaptive meshing algorithm to process the target interaction area. First, the boundary coordinates of the interaction area are obtained, dividing the entire area into uniformly sized square detection grids, with each grid set to a side length of 0.5 meters. The system sets spatial reference markers at the four vertices of each grid for subsequent coordinate transformation and position calibration. Adjacent grids share boundary marker points, forming a continuous grid matrix. The system assigns a unique grid ID to each detection grid and establishes a grid index table for quickly locating and querying motion information within a specific grid.
[0040] Step 302: Collect the user's action video in the target interaction area within a preset time period, and determine the original image sequence based on the action video. The original image sequence includes the user's action process in each detection grid.
[0041] Specifically, high-speed cameras in the target interaction area use a sampling frequency of 60fps to synchronously capture action videos. The field of view of each camera is 120 degrees, and overlapping coverage ensures no-dead-angle capture. The system sets a preset time period of 3 seconds as the action capture window, during which 180 frames of images are continuously captured. The collected video data is processed synchronously by multiple cameras to remove duplicate information and generate an original image sequence with a unified time base. The system preprocesses the original image sequence, including denoising, lighting compensation and distortion correction, to ensure image quality. The processed image sequence clearly records the user's complete action process within each detection grid.
[0042] Step 303: extracting the user's action node positions from the original image sequence, where the action node positions include the spatial coordinates of each human body key point in the detection grid.
[0043] Specifically, a deep learning model analyzes the raw image sequences, identifying and extracting 17 key points of the user's body, including the top of the head, tip of the nose, neck, left and right shoulders, left and right elbows, left and right wrists, center of the torso, left and right hips, left and right knees, and left and right ankles. For each key point, the system calculates its precise coordinates (x, y, z) in three-dimensional space. Through multi-camera view fusion and depth estimation algorithms, sub-centimeter positioning accuracy can be achieved. The spatial coordinates of these key points constitute the position information of the action nodes, which are used for subsequent trajectory analysis.
[0044] Step 304: Determine the moving direction and moving speed of the user action based on the temporal changes of the spatial coordinates.
[0045] Specifically, based on the acquired action node position information, the system analyzes the coordinate changes of each key point between adjacent frames. By calculating the coordinate difference vector, the instantaneous movement direction of each key point is determined. Simultaneously, the instantaneous movement velocity of each key point is calculated based on the inter-frame time interval (1 / 60 second). The system applies a Kalman filter to the velocity data to eliminate sampling noise and produce a smooth velocity curve. By analyzing the changing trends of the velocity curve, the system identifies motion states such as acceleration, constant speed, and deceleration, providing a basis for motion feature extraction.
[0046] Step 305: Generate motion trajectory features according to the moving direction and moving speed.
[0047] Specifically, to accurately generate motion trajectory features that describe the user's complete motion process, the system first analyzes the changes in motion state at adjacent sampling moments. The dot product of the movement direction vectors at every two adjacent sampling moments (t and t+1) is calculated to obtain the direction change angle θ(t). When θ(t) is greater than the preset threshold of 10 degrees, the moment is marked as a potential turning point. The system uses a sliding window (window size of 5 frames) to smooth the angle sequence, eliminating the jitter caused by sampling noise and obtaining a stable direction change curve.
[0048] The system also analyzes velocity changes between adjacent moments. By calculating the velocity difference ΔV(t), it determines the motion state: when ΔV(t) > 0.2 m / s, it is considered an acceleration process; when ΔV(t) < -0.2 m / s, it is considered a deceleration process; and when |ΔV(t)| ≤ 0.2 m / s, it is considered a constant speed process. The system applies a Kalman filter to the velocity data to produce a smooth velocity trend curve.
[0049] Based on the distribution of directional change angles, the system identifies key turning points in an action. When the cumulative angle change exceeds 30 degrees, that point is marked as the demarcation point of an action segment. This allows a continuous action sequence to be broken down into multiple basic action segments, each with a relatively stable motion direction.
[0050] For each action segment, the system extracts the following characteristic parameters: segment duration T, average speed V_avg, maximum speed V_max, speed standard deviation V_std, starting direction D_start, ending direction D_end, cumulative steering angle θ_total, and average angular velocity ω_avg. These parameters comprehensively describe the motion characteristics of the action segment.
[0051] Finally, the feature parameters of all action segments are organized in time sequence to construct a unified feature vector, which is also known as the action trajectory feature. The structure of the feature vector is: [number of segments N, {T1, V_avg1, V_max1, V_std1, D_start1, D_end1, θ_total1, ω_avg1}, ..., {TN, V_avgN, V_maxN, V_stdN, D_startN, D_endN, θ_totalN, ω_avgN}]. This hierarchical feature representation method preserves local action details while reflecting the overall action structure.
[0052] Step 40: Match the motion trajectory features with a preset motion feature library, identify the user's operation intention and generate corresponding control instructions for the target virtual vehicle.
[0053] In the embodiment of the present application, the preset action feature library refers to a data set containing standardized action feature vectors.
[0054] In the embodiment of the present application, the operation intention refers to the specific interactive operation that the system identifies that the user intends to perform on the virtual vehicle.
[0055] Specifically, the system pre-builds an action feature library containing multiple sets of standard action feature templates, each corresponding to a specific operational intent. After acquiring the user's action trajectory features, the system uses a dynamic time warping algorithm to calculate their similarity with each template in the action feature library. The system selects the operational intent corresponding to the template with the highest similarity, exceeding a preset threshold of 0.85. The system then queries a predefined intent-to-command mapping table and converts the identified operational intent into a specific control command for the target virtual vehicle, such as converting "move forward" to "forward (speed=1)." This template-matching-based solution can quickly and accurately identify user intent and generate corresponding control commands.
[0056] Based on the above embodiment, as another optional embodiment, the step of matching the motion trajectory features with a preset motion feature library, identifying the user's operation intention, and generating corresponding control instructions for the target virtual vehicle may further include the following steps: Step 401: extract key action nodes from action trajectory features.
[0057] Specifically, the extreme value analysis method is used to extract key action nodes from the motion trajectory features. First, the velocity curve of each key point of the human body is analyzed to identify the local maximum and minimum points of the velocity. These points usually correspond to the starting, ending or turning positions of the action. The system sets the velocity threshold to 0.5m / s. When the speed change exceeds the threshold, it is marked as a potential key node. At the same time, the acceleration curve is analyzed to identify the inflection points of positive and negative acceleration changes. These inflection points represent the transition of the action state. By fusing the feature points of velocity and acceleration, the system finally determines the timestamps and spatial positions of the key action nodes, providing basic data for the subsequent feature vector construction.
[0058] Step 402: Construct an action feature vector based on the key action nodes. The action feature vector includes spatial features and temporal features of the action.
[0059] Specifically, based on the extracted key action nodes, an action feature vector containing spatial and temporal features is constructed. Spatial features include the normalized coordinates of key points in three-dimensional space, the relative distances between adjacent key points, and the angles between them. Temporal features include the time interval between key nodes, the average velocity, and the rate of change of acceleration. The system organizes these features into fixed-dimensional vectors, with spatial features occupying the first 128 dimensions and temporal features occupying the last 64 dimensions. All eigenvalues are normalized to ensure that the numerical range is between [-1, 1]. This structured feature vector comprehensively describes the spatiotemporal characteristics of the action, facilitating subsequent pattern matching.
[0060] Step 403: Match the action feature vector with the action feature library to obtain a matching result.
[0061] Specifically, a multi-level matching strategy is used to match the action feature vector with the action feature library. First, the locality-sensitive hashing (LSH) algorithm is used to quickly screen out candidate templates in the feature library to reduce the search space. Then, the dynamic time warping (DTW) algorithm is applied to the candidate templates to calculate the similarity score between the feature vector and each template. The system sets the similarity threshold to 0.85 and screens out matching results with scores higher than the threshold. For multiple matching results, the system selects the best match based on a weighted voting mechanism. The weight takes into account the similarity score and historical matching frequency. This progressive matching scheme not only ensures the accuracy of the matching, but also improves processing efficiency.
[0062] Step 404: Determine the user's operation intention based on the matching result, where the operation intention includes the operation direction and operation amplitude.
[0063] Specifically, the system extracts predefined operation type labels from the optimal matching template. It then analyzes the directional component of the feature vector and calculates the unit vector of the primary motion direction to determine the direction of the operation. It also analyzes the displacement of key points to determine the magnitude of the operation. For example, for rotation operations, the angle between the start and end points is used to determine the rotation angle, while for move operations, the modulus of the displacement vector is used to determine the movement distance. The system integrates these parameters into a complete description of the operation intent, including the operation type, direction, and magnitude.
[0064] Step 405: Generate corresponding control instructions for the target virtual vehicle according to the operation intention.
[0065] Specifically, the system pre-establishes a mapping table between operation intentions and control instructions, and generates control instructions in a standard format based on the identified operation intentions. The operation intentions are parsed into basic operation units, such as the rotation operation is parsed into rotation axis, rotation direction, and angle parameters. These parameters are then converted into control instructions in JSON format according to the mapping rules, which contain fields such as instruction type, target object ID, action parameters, etc. For example, the operation of rotating 90 degrees clockwise will generate the instruction: {"type":"rotate","target":"car_001","axis":"y","angle":90,"duration":2000}.
[0066] Based on the above embodiment, as another optional embodiment, the step of determining the user's operation intention based on the matching result may further include the following steps: Step 4041: When the matching result is successful, the operation parameters of the preset action features corresponding to the matching result are used as the user's operation intention.
[0067] Specifically, when the similarity between an action feature vector and a preset action feature in the feature library exceeds 0.85, the system determines that the match is successful. At this point, the system directly extracts the operation parameters from the matching preset action feature record, including the operation type (such as "rotate", "move", "zoom"), operation direction (such as rotation axis vector, movement direction vector), and operation amplitude (such as rotation angle, movement distance, zoom ratio).
[0068] Step 4042: When the matching result is a matching failure, multiple candidate actions are selected from the action feature library, the similarity between the action feature vector and the action feature vector being greater than a preset similarity.
[0069] Specifically, when the highest matching similarity is lower than 0.85, the system activates the candidate action screening mechanism. The system uses the cosine similarity algorithm to calculate the similarity between the action feature vector and all preset actions in the feature library, and selects actions with a similarity greater than 0.7 as the candidate set. To improve retrieval efficiency, the system first uses local sensitive hashing (LSH) pre-screening to block the feature space and perform precise matching only in relevant blocks. The system also considers the spatiotemporal feature similarity and contextual relevance of the action to ensure that the candidate action is consistent with the current interaction scenario.
[0070] Step 4043: Display each candidate action in the target interactive area in the form of an interactive menu, and arrange the candidate actions in the interactive menu from high to low according to similarity.
[0071] Specifically, a translucent arc-shaped interactive menu is generated in the target interaction area, and the menu is located 0.8 meters in front of the user's line of sight. Candidate actions are displayed in the form of icons and text. Each option contains an action type icon, a brief description text and a similarity percentage. The system arranges the candidates in descending order of similarity. The most similar option is located in the center of the menu, and the other options are fanned out. The menu uses a gradient transparency effect to ensure that the user's line of sight to the virtual vehicle is not blocked. Each option is equipped with a visual feedback effect and is highlighted when the user's gaze or gesture hovers over it.
[0072] Step 4044: When a selection operation of the user on the interactive menu is detected within a preset time period, a candidate action corresponding to the selection operation is determined as the user's operation intention.
[0073] Specifically, the system monitors user actions on the interactive menu within a preset 5-second period and supports two selection methods: gesture click and gaze confirmation. Gesture click is triggered by detecting the user's hand hovering over and pressing down on the option area; gaze confirmation requires the user's gaze to remain on the option for at least 0.8 seconds. When a valid selection operation is detected, the system immediately hides the menu and extracts the operation parameters of the selected candidate action as the operation intention. The system also records this selection in the user feedback database to optimize future matching strategies.
[0074] Step 4045: When no user selection operation on the interactive menu is detected within a preset time period, enter the smart assistant mode, which is used to provide interactive operation guidance and action demonstrations.
[0075] Specifically, when no selection operation is detected for more than 5 seconds, the system smoothly transitions to smart assistant mode. First, an anthropomorphic virtual assistant image is displayed in the interactive area, and the assistant uses natural language to introduce the currently available interactive operations. The system also displays action demonstration animations in the virtual space, showing standard action postures through translucent human outlines. The assistant observes the user's actions in real time and provides instant feedback when the user tries to imitate, pointing out the key points of the action and suggestions for improvement. This interactive guidance helps users quickly master the correct operation method. Smart assistant mode lasts until the user successfully completes a valid operation or voluntarily exits.
[0076] Step 50: Transmitting the control instruction to the control center of the target interaction area so that the control center drives the target virtual vehicle.
[0077] In the embodiment of the present application, the control center refers to a local processing unit independently configured in each interactive area, which is responsible for managing the interactive process in the interactive area.
[0078] Specifically, the system allocates an independent communication channel for each interactive zone, using the UDP protocol to transmit control commands in JSON format to the control center of the corresponding target interactive zone. Upon receiving the command, the control center first checks the operating status and resource usage of the virtual vehicle in that zone to ensure that the new command can be executed. The control center then interprets the control command as transformation parameters supported by the zone's rendering engine, and uses the graphics processing unit within that zone to perform the morphological adjustments of the target virtual vehicle. This zone-based distributed control approach ensures independent operation between different interactive zones, avoiding interaction interference.
[0079] Based on the above embodiment, as an optional embodiment, a smart interaction method for an auto show may further include the following process: Specifically, the number of users is detected in real time by using depth cameras deployed in each interactive area. For example, a human detection algorithm based on YOLOv5 can be used to perform image analysis every 100ms to extract the location and ID features of each detected human target. The system uses the DeepSORT algorithm for target tracking and records the entry and exit time of each user. Within the preset statistical period of 60 seconds, the system accumulates the length of stay of each user and adds up the length of stay of all users in the same interactive area to obtain the total length of stay in the area. For example, the system sets the user number threshold to 3 people and the total length of stay threshold to 5 minutes. When it is detected that the number of real-time users in a certain interactive area exceeds 3 people or the total length of stay exceeds 5 minutes, the area is marked as a hot interaction area, triggering the area expansion mechanism.
[0080] Once a hotspot interaction area is identified, the system expands it. Each interaction area is designed with 1.5 times the amount of space available for expansion. This reserved space maintains a safe distance of 1 meter from adjacent interaction areas to ensure expansion does not interfere with each other. The system implements two levels of expansion based on the number of users: 1.2 times the area for three users and 1.5 times for four. If the number of users in an area exceeds four, the system prompts additional users through voice and display prompts to move to other available interaction areas.
[0081] The boundaries of the expanded area are marked with blue laser lines, maintaining a line width of 2cm. The system adjusts the display brightness of the boundary lines based on the number of users: the base brightness is set at 50 nits, increasing to 75 nits for three people, and 100 nits for four people. The boundary lines use a breathing animation effect, with the brightness periodically varying between 80% and 100% of the set value, with a cycle of 2 seconds. Simultaneously, the system synchronously adjusts the interactive devices within the expanded area: recalibrating camera coverage and adjusting the projection parameters of the display devices to ensure the quality of the interactive experience in the expanded space. If the number of users drops below two for a period of three minutes, the system automatically restores the area to its original size. This expansion mechanism, which strictly limits the maximum number of users, ensures the quality of the interactive experience while maintaining the orderly operation of the exhibition area.
[0082] See Figure 2 , is a schematic diagram of a module of an intelligent interactive system for an auto show provided in an embodiment of the present application, wherein the system includes: The area division module is used to divide the target auto show area into multiple independent interactive areas; a vehicle determination module, configured to, when a user triggers an interaction request in a target interaction area, determine a target virtual vehicle based on the interaction request, wherein each interaction area includes a plurality of virtual vehicles; a trajectory determination module, configured to collect a motion video of a user in the target interaction area and determine motion trajectory features of the user based on the motion video; An action matching module is used to match the action trajectory features with a preset action feature library, identify the user's operation intention and generate corresponding control instructions for the target virtual vehicle; The interactive control module is used to transmit the control instruction to the control center of the target interactive area so that the control center drives the target virtual vehicle.
[0083] Optionally, the area division module is further configured to obtain position information of each display vehicle in the target auto show area, wherein the position information includes the center point coordinates and the vehicle body orientation of the display vehicle; Determining a body projection area of each display vehicle based on the position coordinates and body size of each display vehicle; Taking the vehicle body projection area as a reference, an interactive buffer area is expanded outward by a preset distance to obtain the interactive buffer area, wherein the interactive buffer area is used to accommodate user interactive operations; A combined area of the vehicle body projection area and the interactive buffer area is determined as the interactive area corresponding to the display vehicle.
[0084] Optionally, the trajectory determination module is further configured to divide the target interaction area into a plurality of detection grids; Collecting a user's action video within the target interaction area within a preset time period, and determining an original image sequence based on the action video, wherein the original image sequence includes the user's action process within each of the detection grids; Extracting the user's action node positions from the original image sequence, wherein the action node positions include the spatial coordinates of each human body key point in the detection grid; Determining the moving direction and moving speed of the user action based on the temporal changes of the spatial coordinates; The motion trajectory feature is generated according to the moving direction and the moving speed.
[0085] Optionally, the trajectory determination module is further configured to calculate a direction change angle of the user action based on the movement directions at adjacent moments, wherein the direction change angle represents a turning feature of the action; Determine a speed change trend of the user's action based on the movement speed at adjacent moments, wherein the speed change trend includes an acceleration process and a deceleration process; Dividing the user action into multiple action segments according to the distribution of the direction change angles; The motion feature parameters of each of the motion segments are extracted, and the motion trajectory features are generated based on the motion feature parameters of each of the motion segments.
[0086] Optionally, the action matching module is further configured to extract key action nodes from the action trajectory features; Constructing an action feature vector based on the key action node, wherein the action feature vector includes spatial features and temporal features of the action; Matching the action feature vector with the action feature library to obtain a matching result; Determining the user's operation intention based on the matching result, wherein the operation intention includes an operation direction and an operation amplitude; Generate corresponding control instructions for the target virtual vehicle according to the operation intention.
[0087] Optionally, the action matching module is further configured to, when the matching result is successful, use the operation parameters of the preset action feature corresponding to the matching result as the user's operation intention; When the matching result is a matching failure, selecting a plurality of candidate actions from the action feature library whose similarity to the action feature vector is greater than a preset similarity; Displaying each of the candidate actions in the target interaction area in the form of an interactive menu, wherein the candidate actions in the interactive menu are arranged from high to low according to similarity; When a selection operation of the user on the interactive menu is detected within a preset time period, a candidate action corresponding to the selection operation is determined as the user's operation intention; When no user selection operation on the interactive menu is detected within a preset time period, the intelligent assistant mode is entered, and the intelligent assistant mode is used to provide interactive operation guidance and action demonstration.
[0088] Optionally, the intelligent interactive system for an auto show further includes an area adjustment module for detecting the number of users in each interactive area in real time and counting the total residence time of all users in each interactive area within a preset period; When there is a hotspot interaction area where the number of users is greater than the number threshold and / or the total residence time is greater than the time threshold, the hotspot interaction area is expanded according to the number of users and the preset expansion strategy to obtain an extended interaction area. The boundary of the extended interaction area is displayed in a preset color, and the display brightness of the boundary increases as the number of users increases.
[0089] It should be noted that the above embodiments provide systems that implement their functions using only the division of the above functional modules as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0090] An embodiment of the present application further provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded by a processor and executing the intelligent interaction method for a car show in the above embodiment. The specific execution process can be found in the specific description of the above embodiment and will not be repeated here.
[0091] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 The electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .
[0092] The communication bus 302 is used to implement the connection and communication between these components.
[0093] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0094] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0095] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.
[0096] Among them, the memory 305 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also optionally be at least one storage device located away from the aforementioned processor 301. Refer to Figure 3 , as a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface module, and an application program of an intelligent interaction method for an auto show.
[0097] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call an application program stored in the memory 305 for a location guidance method for blind people to travel. When executed by one or more processors 301, the electronic device 300 executes one or more methods in the above-mentioned embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0098] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0100] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0101] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0103] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.
[0104] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.
Claims
1. An intelligent interactive method for an auto show, characterized in that: The method comprises: Divide the target auto show area into multiple independent interactive areas; When a user triggers an interaction request in a target interaction area, determining a target virtual vehicle based on the interaction request, wherein each interaction area includes a plurality of virtual vehicles; Collecting a motion video of a user in the target interaction area, and determining a motion trajectory feature of the user based on the motion video; Matching the motion trajectory features with a preset motion feature library, identifying the user's operation intention and generating corresponding control instructions for the target virtual vehicle; The control instruction is transmitted to a control center of the target interaction area, so that the control center drives the target virtual vehicle.
2. The intelligent interaction method for an auto show according to claim 1, characterized in that: The target auto show area is divided into multiple independent interactive areas, including: Obtaining the location information of each display vehicle in the target auto show area, wherein the location information includes the center coordinates and body orientation of the display vehicle; Determining a body projection area of each display vehicle based on the position information and body size of each display vehicle; Taking the vehicle body projection area as a reference, an interactive buffer area is expanded outward by a preset distance to obtain the interactive buffer area, wherein the interactive buffer area is used to accommodate user interactive operations; A combined area of the vehicle body projection area and the interactive buffer area is determined as the interactive area corresponding to the display vehicle.
3. The intelligent interaction method for an auto show according to claim 1, characterized in that: The collecting of the action video of the user in the target interaction area and determining the action trajectory characteristics of the user based on the action video includes: Dividing the target interaction area into a plurality of detection grids; Collecting a user's action video within the target interaction area within a preset time period, and determining an original image sequence based on the action video, wherein the original image sequence includes the user's action process within each of the detection grids; Extracting the user's action node positions from the original image sequence, wherein the action node positions include the spatial coordinates of each human body key point in the detection grid; Determining the moving direction and moving speed of the user action based on the temporal changes of the spatial coordinates; The motion trajectory feature is generated according to the moving direction and the moving speed.
4. The intelligent interaction method for an auto show according to claim 3, characterized in that: The generating the motion trajectory feature according to the moving direction and the moving speed includes: Calculating a direction change angle of the user's action based on the movement directions at adjacent moments, wherein the direction change angle represents a turning feature of the action; Determine a speed change trend of the user's action based on the movement speed at adjacent moments, wherein the speed change trend includes an acceleration process and a deceleration process; Dividing the user action into multiple action segments according to the distribution of the direction change angle and the speed change trend; The motion feature parameters of each of the motion segments are extracted, and the motion trajectory features are generated based on the motion feature parameters of each of the motion segments.
5. The intelligent interaction method for an auto show according to claim 1, characterized in that: The matching of the motion trajectory feature with a preset motion feature library, identifying the user's operation intention and generating a corresponding control instruction for the target virtual vehicle includes: Extracting key action nodes from the action trajectory features; Constructing an action feature vector based on the key action node, wherein the action feature vector includes spatial features and temporal features of the action; Matching the action feature vector with the action feature library to obtain a matching result; Determining the user's operation intention based on the matching result, wherein the operation intention includes an operation direction and an operation amplitude; Generate corresponding control instructions for the target virtual vehicle according to the operation intention.
6. The intelligent interaction method for an auto show according to claim 5, characterized in that: Determining the user's operation intention based on the matching result includes: When the matching result is successful, the operation parameter of the preset action feature corresponding to the matching result is used as the user's operation intention; When the matching result is a matching failure, selecting a plurality of candidate actions from the action feature library whose similarity to the action feature vector is greater than a preset similarity; Displaying each of the candidate actions in the target interaction area in the form of an interactive menu, wherein the candidate actions in the interactive menu are arranged from high to low according to similarity; When a selection operation of the user on the interactive menu is detected within a preset time period, a candidate action corresponding to the selection operation is determined as the user's operation intention; When no user selection operation on the interactive menu is detected within a preset time period, the intelligent assistant mode is entered, and the intelligent assistant mode is used to provide interactive operation guidance and action demonstration.
7. The intelligent interaction method for an auto show according to claim 1, characterized in that: The method further comprises: Detecting the number of users in each interactive area in real time, and counting the total residence time of all users in each interactive area within a preset period; When there is a hotspot interaction area where the number of users is greater than the number threshold and / or the total residence time is greater than the time threshold, the hotspot interaction area is expanded according to the number of users and the preset expansion strategy to obtain an extended interaction area. The boundary of the extended interaction area is displayed in a preset color, and the display brightness of the boundary increases as the number of users increases.
8. An intelligent interactive system for an auto show, characterized in that: The system comprises: The area division module is used to divide the target auto show area into multiple independent interactive areas; a vehicle determination module, configured to, when a user triggers an interaction request in a target interaction area, determine a target virtual vehicle based on the interaction request, wherein each interaction area includes a plurality of virtual vehicles; a trajectory determination module, configured to collect a motion video of a user in the target interaction area and determine motion trajectory features of the user based on the motion video; An action matching module is used to match the action trajectory features with a preset action feature library, identify the user's operation intention and generate corresponding control instructions for the target virtual vehicle; The interactive control module is used to transmit the control instruction to the control center of the target interactive area so that the control center drives the target virtual vehicle.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises a processor, a memory, a user interface and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Providing an interactive experience using a 3D depth camera and a 3D projector
CN102693005A
Intel Real Sense-based automobile display application development method and system
CN107861714A
Space-based situation simulation system based on mixed reality technology
CN114612640A
Virtual interaction method and device, equipment and medium
CN116614543A
Vehicle exhibition implementation method, system and equipment based on cloud exhibition and medium
CN116931732A
Cited By
Virtual digital image interaction method and system for precise advertisement playing
CN121010740A
Conference content real-time screen projection and interactive labeling method of ultra-high-definition display system
CN121486525A