Vehicle-mounted image rendering method and device based on YTS engine AI algorithm
By using the YTS engine AI algorithm to identify and strengthen the rendering of key action areas in car image rendering, the user experience problem caused by excessive rendering of car image is solved, and more efficient user attention guidance and important information recognition are achieved.
Patent Information
- Application Number
- CN202510254777.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing on-board image rendering content is too much and the dynamic transformation frequency is high, making it difficult for on-board users to pay attention to key content in the image, resulting in poor user experience.
The vehicle image rendering method based on the YTS engine AI algorithm is adopted to identify the main objects and dynamic actions in the vehicle video, determine the action picture position range, and enhance the rendering of the important focus position range when the vehicle video data is played, guiding users to pay attention to key content.
Effectively guide vehicle users to pay attention to important content during the playback of vehicle videos, improve users' awareness and understanding of key events or actions, and improve vehicle user experience.
Smart Images

Figure CN119762651B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of rendering technology, and in particular to a vehicle-mounted image rendering method and device based on the YTS engine AI algorithm. Background Art
[0002] At present, in-vehicle image rendering technology refers to the use of computer graphics and related algorithms to process and optimize image or video content in real time in the automotive environment to meet the special requirements of in-vehicle display devices, such as instrument panels, central control screens, head-up displays (HUDs), etc. This technology has gradually emerged with the development of intelligent and connected vehicles, and plays a vital role in improving driving safety and user experience.
[0003] However, the existing vehicle-mounted image rendering has a large amount of content and a high dynamic change frequency, which makes it difficult for vehicle-mounted users to focus on the most important key content in the vehicle-mounted image, resulting in a poor experience for vehicle-mounted users due to the vehicle-mounted image rendering effect. Summary of the invention
[0004] The purpose of the present invention is to provide a vehicle-mounted image rendering method and device based on the YTS engine AI algorithm to solve the technical problem that the vehicle-mounted image rendering effect causes poor vehicle-mounted user experience.
[0005] In a first aspect, the present application provides a vehicle-mounted image rendering method based on the YTS engine AI algorithm, the method comprising:
[0006] Obtain vehicle video data;
[0007] Identify a first subject object in the vehicle-mounted video data by using an AI algorithm and analyze a dynamic action corresponding to the first subject object in the vehicle-mounted video data;
[0008] Determining a target frame containing the execution process corresponding to the dynamic action from the multiple frames corresponding to the vehicle-mounted video data;
[0009] For each target frame in the plurality of target frames, determining an action frame position range corresponding to the dynamic action performed by the first subject object in the target frame;
[0010] Determine the next second target frame of each first target frame according to the sequential frame playback sequence of the target frame in the vehicle video data, and determine the first action picture position range corresponding to the first target frame and the second action picture position range corresponding to the second target frame based on the action picture position range corresponding to each target frame; wherein the first target frame is any target frame in the multiple frames corresponding to the vehicle video data that contains the corresponding execution process of the dynamic action;
[0011] Under the condition that the relative positions of the first target frame picture and the second target frame picture are kept unchanged, determining a third action picture position range corresponding to the second action picture position range in the first target frame picture, and determining the first action picture position range and the third action picture position range as important focus position ranges in the first target frame picture;
[0012] When the vehicle-mounted video data is played to the first target frame, the important focus position range is enhanced and rendered.
[0013] In a possible implementation, after determining the first action picture position range and the third action picture position range as the important focus position range in the first target frame picture, the method further includes:
[0014] Determine the screen distance between each position in the first target frame and the important focus position range;
[0015] Determining the weakened rendering degree of each position according to the screen distance; the farther the screen distance is, the greater the weakened rendering degree is;
[0016] According to the weakening rendering degree of each position, the picture position range other than the important focus position range in the first target frame picture is weakened rendering according to the corresponding weakening rendering degree, so as to guide the in-vehicle user to focus on the important focus position range and the peripheral position range of the important focus position range during the playback of the first target frame picture.
[0017] In a possible implementation, after determining the first action picture position range and the third action picture position range as the important focus position range in the first target frame picture, the method further includes:
[0018] When the in-vehicle video data is played to the first target frame picture, the important focus position range is rendered according to the first specified rendering method, and the picture position range other than the important focus position range in the first target frame picture is rendered according to the second specified rendering method, so as to guide the in-vehicle user to focus on the first action picture position range during the playback of the first target frame picture and to focus on the second action picture position range in advance through the third action picture position range during the playback from the first target frame picture to the second target frame picture;
[0019] Among them, the first specified rendering method includes at least one of ray tracing rendering, color rendering and animation style rendering; the second specified rendering method includes black and white rendering and a rendering method that keeps the picture style unchanged.
[0020] In a possible implementation, rendering the important focus position range according to a first specified rendering mode, and rendering a picture position range other than the important focus position range in the first target frame picture according to a second specified rendering mode, includes:
[0021] Decomposing the composite rendering task to be rendered into a plurality of sub-rendering tasks, and setting the plurality of sub-rendering tasks in a task queue; wherein the plurality of sub-rendering tasks include a first sub-rendering task for rendering the important focus position range according to a first specified rendering method, and a second sub-rendering task for rendering a screen position range other than the important focus position range in the first target frame screen according to a second specified rendering method;
[0022] The plurality of sub-rendering tasks are dynamically allocated to a plurality of threads according to the priorities and dependencies of the tasks, and the plurality of sub-rendering tasks are executed by the plurality of threads.
[0023] In a possible implementation, after determining the target frame image containing the execution process corresponding to the dynamic action from the multiple frames corresponding to the vehicle-mounted video data, the method further includes:
[0024] Performing image recognition on the target frame to obtain an image recognition result, and judging whether the target first subject object in the target frame conforms to a specified logic rule according to the image recognition result;
[0025] If the target first main object does not conform to the specified logical rule, the target first main object is eliminated to obtain a target frame picture after the elimination process.
[0026] In a possible implementation, after acquiring the vehicle-mounted video data, the method further includes:
[0027] Acquire film and television resource image data from the vehicle-mounted video data, and identify a second subject object in the film and television resource image data;
[0028] Analyze the film and television resource image data and the style data of the second subject object through the YTS engine;
[0029] For the second main object in a plurality of different specified scenes, generating object appearance data corresponding to the second main object according to the style data;
[0030] The second main object is rendered based on the object appearance data to display the second main object in accordance with the style corresponding to the specified scene.
[0031] In a possible implementation, the first subject object includes any one or more of the following:
[0032] Vehicle objects, character objects, and static object objects.
[0033] In a second aspect, the present application provides a vehicle-mounted image rendering device based on the YTS engine AI algorithm, comprising:
[0034] An acquisition module, used to acquire vehicle-mounted video data;
[0035] An identification module, configured to identify a first subject object in the vehicle-mounted video data by using an AI algorithm and analyze a dynamic action corresponding to the first subject object in the vehicle-mounted video data;
[0036] A first determination module is used to determine a target frame containing the execution process corresponding to the dynamic action from multiple frames corresponding to the vehicle-mounted video data;
[0037] A second determining module is used to determine, for each target frame in the plurality of target frames, an action picture position range corresponding to the dynamic action performed by the first subject object in the target frame;
[0038] A third determination module is used to determine the next second target frame of each first target frame according to the sequential frame playback sequence of the target frames played in the vehicle-mounted video data, and determine the first action picture position range corresponding to the first target frame and the second action picture position range corresponding to the second target frame based on the action picture position range corresponding to each target frame; wherein the first target frame is any target frame in the multiple frames corresponding to the vehicle-mounted video data that contains the corresponding execution process of the dynamic action;
[0039] a fourth determination module, configured to determine a third action picture position range corresponding to the second action picture position range in the first target frame picture while keeping the relative overall picture positions of the first target frame picture and the second target frame picture unchanged, and determine the first action picture position range and the third action picture position range as important focus position ranges in the first target frame picture;
[0040] A rendering module is used to enhance the rendering of the important focus position range when the vehicle-mounted video data is played to the first target frame.
[0041] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method described in the first aspect.
[0043] This application brings the following beneficial effects:
[0044] The present application provides a vehicle-mounted image rendering method and device based on the YTS engine AI algorithm, which can obtain vehicle-mounted video data, identify the first main object in the vehicle-mounted video data through the AI algorithm and analyze the dynamic action corresponding to the first main object in the vehicle-mounted video data, determine the target frame picture containing the corresponding execution process of the dynamic action from the multiple frames corresponding to the vehicle-mounted video data, determine the action picture position range corresponding to the dynamic action performed by the first main object in the target frame picture for each frame of the multiple frames of the target frame pictures, determine the next frame of the second target frame picture of each frame of the first target frame picture according to the sequential frame playback sequence of the multiple frames of the target frame pictures played in the vehicle-mounted video data, and determine the first action picture position range corresponding to the first target frame picture and the second action picture position range corresponding to the second target frame picture based on the action picture position range corresponding to each frame of the target frame picture, wherein the first target frame picture is any frame of the target frame picture containing the corresponding execution process of the dynamic action in the multiple frames of the target frame pictures corresponding to the vehicle-mounted video data, while keeping the relative overall picture position between the first target frame picture and the second target frame picture unchanged. , determine the third action picture position range corresponding to the second action picture position range in the first target frame picture, and determine the first action picture position range and the third action picture position range as the important focus position range in the first target frame picture, and perform enhanced rendering on the important focus position range when the in-vehicle video data is played to the first target frame picture, so as to guide the in-vehicle user to focus on the first action picture position range during the playback of the first target frame picture and to pay attention to the second action picture position range in advance through the enhanced rendering of the third action picture position range during the playback from the first target frame picture to the second target frame picture, that is, to enhance the rendering of the third action picture position range in advance in the transition from the first target frame to the second target frame, which can help the user to notice the upcoming changes in advance, achieve a predictive prompt effect, and reduce cognitive load. Moreover, by enhancing the rendering of a specific action picture position range, the user's visual attention can be effectively guided to key information, thereby guiding the in-vehicle user's attention, improving the user's cognition and understanding of important events or actions, and improving the in-vehicle user experience, thereby solving the technical problem that the in-vehicle image rendering effect causes the in-vehicle user experience to be poor.
[0045] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 A schematic diagram of a process flow of a vehicle-mounted image rendering method based on the YTS engine AI algorithm provided in an embodiment of the present application;
[0048] Figure 2 Another schematic diagram of the process of the vehicle-mounted image rendering method based on the YTS engine AI algorithm provided in an embodiment of the present application;
[0049] Figure 3 A schematic diagram of the structure of a vehicle-mounted image rendering device based on the YTS engine AI algorithm provided in an embodiment of the present application;
[0050] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.
[0052] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.
[0053] At present, the existing in-vehicle image rendering content is large and the dynamic change frequency is fast, which makes it difficult for in-vehicle users to pay attention to the most important key content in the in-vehicle image, resulting in a poor in-vehicle user experience due to the in-vehicle image rendering effect. Based on this, the embodiment of the present application provides an in-vehicle image rendering method and device based on the YTS engine AI algorithm, which can solve the technical problem that the in-vehicle image rendering effect causes a poor in-vehicle user experience.
[0054] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.
[0055] Figure 1 The flowchart of a vehicle image rendering method based on the YTS engine AI algorithm provided in the embodiment of the present application is as follows. Figure 1 As shown, the method includes:
[0056] Step S110, obtaining vehicle-mounted video data.
[0057] As a possible implementation, video capture is first performed, that is, the analog signal captured by the camera is converted into a digital video stream using a hardware encoder or a software encoder. The digital video stream can be uncompressed or compressed through standards such as H.264 and H.265 to save bandwidth and storage space. Then, the video data is transmitted to the central processing unit or directly stored through the vehicle network (such as CAN bus, MOST, Ethernet, etc.). For situations where real-time viewing is required, it may also be necessary to send the data to a remote server through a wireless communication module (such as 4G / 5G, Wi-Fi). Afterwards, the video data may need to be pre-processed, such as denoising, stabilization, etc. Computer vision algorithms are applied to analyze the video stream, such as target detection, lane departure warning, pedestrian recognition, etc. Then, the video data is stored in a local storage device (such as SSD, HDD) for subsequent review or evidence preservation.
[0058] It should be noted that the YTS (Unity TV Service) engine in the embodiment of the present application represents the Unity visual rendering service engine. Unity is a real-time 3D interactive content creation and operation platform. All creators, including game development, art, architecture, car design, and film and television, use Unity to turn their ideas into reality. The platform provides a complete set of software solutions that can be used to create, operate, and realize any real-time interactive 2D and 3D content. Supported platforms include mobile phones, tablets, PCs, game consoles, augmented reality, and virtual reality devices.
[0059] Step S120: identifying the first main object in the vehicle-mounted video data through an AI algorithm and analyzing the dynamic actions corresponding to the first main object in the vehicle-mounted video data.
[0060] The first main object includes any one or more of the following: a vehicle object, a person object, and a static object object.
[0061] Exemplarily, target detection and tracking are performed first, that is, pre-trained deep learning models (such as YOLO, SSD, Faster R-CNN, etc.) are used to detect the main object in the video frame. Once the main object is detected, it is continuously tracked using a multi-target tracking algorithm (such as SORT, Deep SORT), and can be re-identified even after a short occlusion. Authentication is then performed. For specific applications (such as driver monitoring), additional authentication steps may be required, such as facial recognition technology to confirm the identity of the first main object. Behavior analysis is then performed, that is, the behavior pattern of the main object is analyzed using time series data, which can be achieved through models such as RNN (recurrent neural network) and LSTM (long short-term memory network). The dynamic movements of the main object are analyzed, and based on these behaviors, it is determined whether there are safety hazards or other important events. In terms of result output and feedback, the analysis results are presented to the user in a visual manner, such as displaying a real-time annotated video stream on a display.
[0062] Step S130 , determining a target frame image including a corresponding execution process of a dynamic action from the multiple frames of images corresponding to the vehicle-mounted video data.
[0063] For example, first perform inter-frame difference analysis, that is, calculate the difference between adjacent frames to identify areas that may contain dynamic changes. This can be achieved by calculating the pixel-level difference between two frames or using the optical flow method. Mark frames with significant changes, which may be key points where the subject starts or ends an action. Then, perform target detection and tracking initialization, apply target detection algorithms (such as Faster R-CNN, YOLO, etc.) on the selected frames, identify and select the first subject. Initialize the tracker (such as KCF, Deep SORT, etc.) to continuously track the position of the subject in subsequent frames. Then perform key point detection and pose estimation. For the marked change frames, apply the human pose estimation model (such as OpenPose, MediaPipe, etc.) to detect the key points of the subject. Analyze the position and changes of these key points to understand the action state of the subject. Then, for behavioral pattern recognition, use machine learning or deep learning models (such as LSTM, GRU, etc.) to analyze time series data, that is, pose changes over multiple consecutive frames to identify specific behavioral patterns. By comparing with a pre-defined behavior pattern library, it is determined which frames contain the start, middle or end stages of the dynamic action of interest.
[0064] In terms of target frame selection, based on the results of behavioral pattern recognition, a frame sequence that can fully display the entire process of a dynamic action is selected. If necessary, key frames can be further filtered based on the importance of the action or user-defined rules. In terms of result output and visualization, the selected target frame and its related information (such as timestamp, action type, etc.) are output to the user interface or storage system.
[0065] Step S140 : for each target frame in the plurality of target frames, determining an action frame position range corresponding to the dynamic action performed by the first subject object in the target frame.
[0066] As an optional implementation, perform subject object detection and positioning first, and apply target detection algorithms (such as Faster R-CNN, YOLO, etc.) in each frame to accurately locate and select the first subject object. Output the bounding box coordinates (x_min, y_min, x_max, y_max) of the subject object, which define the position range of the subject object in the current frame. Then, perform key point detection and pose estimation, and apply human pose estimation models (such as OpenPose, MediaPipePose, etc.) to detect the key parts of the subject object (such as joint positions). Obtain the two-dimensional coordinates of each key point, and infer the pose and action mode of the subject object based on these coordinates.
[0067] For the recognition and classification of dynamic actions, a pre-trained action recognition model (such as a model based on LSTM or 3D CNN) is used to analyze the posture changes on consecutive frames and identify the specific dynamic action type. Based on the recognized action type, context information is provided for subsequent steps, which helps to more accurately define the position range of the action screen.
[0068] For the determination of the action screen position range, the screen area occupied by the subject object when performing a specific action is calculated based on the position information of the bounding box and key points of the subject object. If the action involves a large range of motion (such as jumping or waving), it may be necessary to consider the changes of the subject object between frames and the maximum value of its projected area. For some complex actions, it may also be necessary to consider the directionality and speed of the action in order to more accurately define the area affected by the action. After that, a spatiotemporal consistency check can be performed, that is, to check whether the position and posture of the subject object between adjacent frames are coherent to ensure that the determined action screen position range conforms to physical logic and time sequence. Correct misjudgments caused by rapid movement or other factors to ensure the accuracy of the action screen position range. For the output of the results, the action screen position range corresponding to the dynamic action performed by the subject object in each frame is finally determined and marked with a rectangular frame or other forms.
[0069] Step S150, determining the next frame of the second target frame of each frame of the first target frame according to the sequential frame playback sequence of the multiple frames of the target frame in the vehicle video data, and determining the corresponding first action picture position range in the first target frame and the corresponding second action picture position range in the second target frame based on the corresponding action picture position range in each frame of the target frame.
[0070] The first target frame is any target frame that contains a corresponding execution process of a dynamic action among the multiple frames corresponding to the vehicle-mounted video data.
[0071] For example, the frame playback sequence is first established. For parsing the video stream, the vehicle video data is first parsed, all frames are extracted, and sorted according to the original recording timestamps to ensure the correct playback order. For marking key frames, key frames containing dynamic actions are identified in all frames. This can be achieved through a pre-trained action detection model that can identify the start and end of specific dynamic actions.
[0072] Then, the first target frame and its next frame (i.e., the second target frame) are determined, and the first target frame is selected: any frame is selected from the above-marked key frames as the first target frame. The second target frame is located: according to the frame playback sequence, a frame after the first target frame is directly obtained as the second target frame. Then, the action frame position range is defined. For the first target frame: subject object detection and positioning: the target detection algorithm is applied to accurately locate the position of the first subject object. Pose estimation: the key point coordinates of the subject object are obtained through the human body pose estimation or object pose estimation model. Determine the first action frame position range: combine the subject object boundary box and key point information to define the action frame position range corresponding to the dynamic action performed by the subject object in the first target frame. For the second target frame: repeat subject object detection and positioning: similarly, locate the same subject object in the second target frame. Update pose estimation: use the pose estimation model again to update the pose information of the subject object. Determine the second action frame position range: according to the updated subject object position and pose, define the action frame position range corresponding to the dynamic action performed by the subject object in the second target frame.
[0073] Then, the comparison and adjustment process is performed to compare the changes between the two frames: compare the position and posture differences of the main object between the first target frame and the second target frame to evaluate the development of dynamic actions. Check spatiotemporal consistency: ensure the continuity of the action between the two frames and correct possible misjudgments caused by rapid movement or other factors. Optimize the position range: if necessary, adjust the position range of the action picture according to the comparison results between the two frames to ensure its accuracy and rationality.
[0074] As for the result output, the position range of the action pictures in the first target frame and the second target frame are marked with a rectangular frame or other forms, and relevant information such as action type, frame number, etc. are output. Subsequent processing: This information can be used for further analysis, such as behavior monitoring, anomaly detection, or the development of assisted driving systems.
[0075] Step S160, while keeping the relative overall picture position between the first target frame picture and the second target frame picture unchanged, determine the third action picture position range corresponding to the second action picture position range in the first target frame picture, and determine the first action picture position range and the third action picture position range as important focus position ranges in the first target frame picture.
[0076] As a possible implementation method, a reference coordinate system is determined and a reference frame is selected: the first target frame is selected as the reference frame, and all subsequent calculations will be performed based on this frame. A coordinate system is established: a coordinate system is defined in the reference frame for accurately locating elements in the picture. Then, the position of the action picture is identified and feature extraction is performed: feature points are extracted from the first target frame and the second target frame using computer vision algorithms (such as SIFT, SURF, ORB, etc.). Feature points are matched: the corresponding relationship is found by comparing the feature points between the two frames, which helps to keep the relative overall picture position between the two unchanged. Action area is calculated: the position range of the first action picture in the first target frame and the position range of the second action picture in the second target frame are determined based on the matched feature points. After that, the transformation matrix is calculated and the transformation parameters are estimated: the RANSAC algorithm or a similar method is used to estimate the geometric transformation (translation, rotation, scaling, etc.) from the second target frame to the first target frame to ensure that the relative position is consistent even if the camera moves. Construct a transformation matrix: construct a transformation matrix based on the estimated transformation parameters.
[0077] Then, map the second action picture position and apply the transformation matrix: input the data points of the second action picture position range into the transformation matrix and calculate their new positions in the first target frame, i.e., the third action picture position range. Adjust the bounding box: If necessary, perform bounding box adjustment on the mapped action picture position to fit the aspect ratio and size of the first target frame.
[0078] Then determine the important focus position range. Merge area: Perform a logical operation (such as taking a union) on the first action picture position range and the third action picture position range to obtain the important focus position range in the final first target frame. Optimize area: It may be necessary to do some post-processing on the merged area, such as removing redundant small areas or connecting adjacent areas, to improve the quality of the result.
[0079] Step S170 , when the vehicle-mounted video data is played to the first target frame, an important focus position range is enhanced and rendered.
[0080] In actual applications, determine how and when to apply visual enhancement effects based on demand, such as color and brightness changes, adding graphic elements (arrows, frames, etc.), or animation effects. Dynamic adjustment: Consider the impact of the vehicle's driving status (speed, steering, etc.) on user attention and adjust the rendering intensity or method in a timely manner.
[0081] For real-time processing and rendering, real-time detection and rendering: When the video playback reaches the first target frame, immediately start the enhanced rendering of the important focus position range. This process requires low latency to ensure user experience. Smooth transition: Design a reasonable transition effect to make the switch from the first target frame to the second target frame natural and smooth, avoiding discomfort caused by sudden changes.
[0082] For the user-guided mechanism, advance prompts can be performed: before approaching the second target frame, start to slightly enhance the rendering of the third action frame position range, gradually increase its prominence until it completely enters the field of vision. Continuous reminders: maintain a certain level of visual prompts throughout the transition process to help users shift their eyes to the new focus that is about to appear.
[0083] In terms of in-car user feedback, we can understand the user's actual viewing behavior through eye tracking or other interactive methods, and evaluate the effect of enhanced rendering accordingly. In terms of iterative improvements, we can continuously optimize algorithm parameters and rendering strategies based on the collected data to improve the effectiveness of the system and user experience.
[0084] In an embodiment of the present application, by performing enhanced rendering on an important focus position range when the in-vehicle video data is played to the first target frame, the in-vehicle user can be guided to focus on the first action picture position range during the playback of the first target frame and to pay attention to the second action picture position range in advance through the enhanced rendering of the third action picture position range during the playback from the first target frame to the second target frame. Specifically, by enhancing the rendering of the third action picture position range in advance during the transition from the first target frame to the second target frame, the user can notice the upcoming changes in advance, achieve a predictive prompt effect, and reduce cognitive load. Moreover, by performing enhanced rendering on a specific action picture position range, the user's visual attention can be effectively guided to key information, thereby guiding the in-vehicle user's attention, improving the user's cognition and understanding of important events or actions, and enhancing the in-vehicle user experience.
[0085] In some embodiments, after the above step S160, the method may further include the following steps:
[0086] Determine the screen distance between each position in the first target frame and the important focus position range; determine the weakening rendering degree of each position according to the screen distance; the farther the screen distance is, the greater the weakening rendering degree is;
[0087] According to the weakening rendering degree of each position, the picture position range other than the important focus position range in the first target frame picture is weakened rendering according to the corresponding weakening rendering degree, so as to guide the in-vehicle user to focus on the important focus position range and the peripheral position range of the important focus position range during the playback of the first target frame picture.
[0088] By weakening the image of non-critical areas, the user's visual focus is more easily concentrated on the important focus position range and its surroundings, focusing on key information, thereby improving the cognitive efficiency of key information. Furthermore, the weakening process reduces the attraction of the surrounding environment to the user's sight, avoids unnecessary distractions, reduces interference, and helps keep the driver's attention focused on important dynamics.
[0089] In some embodiments, after the above step S160, the method may further include the following steps:
[0090] When the in-vehicle video data is played to the first target frame picture, the important focus position range is rendered according to the first specified rendering method, and the picture position range other than the important focus position range in the first target frame picture is rendered according to the second specified rendering method, so as to guide the in-vehicle user to focus on the first action picture position range during the playback of the first target frame picture and to focus on the second action picture position range in advance through the third action picture position range during the playback from the first target frame picture to the second target frame picture;
[0091] Among them, the first specified rendering method includes at least one of ray tracing rendering, color rendering and animation style rendering; the second specified rendering method includes black and white rendering and a rendering method with unchanged picture style.
[0092] By using the first designated rendering method (such as ray tracing, color or animation style rendering) for the important focus position range, these areas are made more prominent visually, effectively attracting the user's attention and achieving the effect of optimizing visual guidance. Furthermore, by using the second designated rendering method (black and white rendering or keeping the original style unchanged) for other parts of the screen, non-critical information is weakened, reducing the interference of non-critical information, allowing users to focus more on important visual elements.
[0093] In some embodiments, the above-mentioned rendering of the important focus position range according to the first specified rendering mode, and rendering of the screen position range other than the important focus position range in the first target frame screen according to the second specified rendering mode, may specifically include the following steps:
[0094] Decomposing the composite rendering task to be rendered into a plurality of sub-rendering tasks, and setting the plurality of sub-rendering tasks in a task queue; wherein the plurality of sub-rendering tasks include a first sub-rendering task for rendering an important focus position range according to a first specified rendering method, and a second sub-rendering task for rendering a picture position range other than the important focus position range in a first target frame picture according to a second specified rendering method;
[0095] According to the priorities and dependencies of the tasks, multiple sub-rendering tasks are dynamically assigned to multiple threads, and the multiple sub-rendering tasks are executed through multiple threads.
[0096] By breaking down the composite rendering task into multiple subtasks and executing them concurrently using multiple threads, the time required for the entire rendering process can be significantly shortened. Each thread focuses on a specific type of rendering work (such as the first specified rendering mode or the second specified rendering mode), which makes more efficient use of resources. Moreover, tasks are dynamically assigned to different threads based on task priorities and dependencies, ensuring that tasks on the critical path can be completed as quickly as possible while also avoiding delays caused by waiting for other non-critical tasks.
[0097] In some embodiments, after the above step S130, the method may further include the following steps:
[0098] Performing image recognition on the target frame to obtain an image recognition result, and judging whether the target first subject object in the target frame conforms to a specified logical rule according to the image recognition result;
[0099] If the target first main object does not conform to the specified logical rule, the target first main object is eliminated to obtain a target frame picture after the elimination process.
[0100] By performing image recognition on the main objects in each frame and judging whether these objects meet the requirements according to preset logical rules, the content of all frames in the video or image sequence is guaranteed to remain consistent and coherent. This is very important for application scenarios that need to maintain specific visual specifications, such as advertising and educational video production.
[0101] In some embodiments, after the above step S110, Figure 2 As shown, the method may also include the following steps:
[0102] Step S210, acquiring film and television resource image data from the vehicle-mounted video data, and identifying a second subject object in the film and television resource image data;
[0103] Step S220, analyzing the film and television resource image data and the style data of the second main object through the YTS engine;
[0104] Step S230, for the second main object in a plurality of different specified scenes, generating object appearance data corresponding to the second main object according to the style data;
[0105] Step S240: Render the second main object based on the object appearance data to display the second main object in a style corresponding to the specified scene.
[0106] By stylizing the second subject objects in different specified scenes, the appearance of the objects finally presented to the user can better match the user's preferences or the needs of the current environment. For example, when driving at night, the system may automatically adjust the color and brightness of the display elements in the car to reduce the impact on the driver's line of sight. Furthermore, the YTS engine is used to analyze the image data of film and television resources and the style data of the second subject objects, and based on this information, the object appearance data that conforms to the specific scene is generated, thereby ensuring that the rendered objects are more visually realistic and natural, enhancing the realism of the overall picture and the user's immersive experience. This is particularly important for entertainment applications (such as in-car movie playback), educational simulation training, etc.
[0107] Figure 3 A schematic diagram of the structure of a vehicle-mounted image rendering device based on the YTS engine AI algorithm is provided. Figure 3 As shown, the vehicle-mounted image rendering device 300 based on the YTS engine AI algorithm includes:
[0108] An acquisition module 301 is used to acquire vehicle-mounted video data;
[0109] An identification module 302, configured to identify a first subject object in the vehicle-mounted video data by using an AI algorithm and analyze a dynamic action corresponding to the first subject object in the vehicle-mounted video data;
[0110] A first determination module 303 is used to determine a target frame containing the execution process corresponding to the dynamic action from multiple frames corresponding to the vehicle-mounted video data;
[0111] A second determination module 304 is used to determine, for each target frame in the plurality of target frames, an action frame position range corresponding to the dynamic action performed by the first subject object in the target frame;
[0112] The third determination module 305 is used to determine the next second target frame of each first target frame according to the sequential frame playback sequence of the multiple frames of the target frame played in the vehicle-mounted video data, and determine the first action picture position range corresponding to the first target frame and the second action picture position range corresponding to the second target frame based on the action picture position range corresponding to each frame of the target frame; wherein the first target frame is any target frame in the multiple frames corresponding to the vehicle-mounted video data that contains the corresponding execution process of the dynamic action;
[0113] A fourth determination module 306 is used to determine a third action picture position range corresponding to the second action picture position range in the first target frame picture while keeping the relative overall picture positions of the first target frame picture and the second target frame picture unchanged, and determine the first action picture position range and the third action picture position range as important focus position ranges in the first target frame picture;
[0114] The rendering module 307 is used to enhance the rendering of the important focus position range when the vehicle-mounted video data is played to the first target frame.
[0115] The vehicle-mounted image rendering device based on the YTS engine AI algorithm provided in the embodiment of the present application has the same technical features as the vehicle-mounted image rendering method based on the YTS engine AI algorithm provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.
[0116] An electronic device provided in an embodiment of the present application is Figure 4 As shown, the electronic device 400 includes a processor 402 and a memory 401, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps of the method provided in the above embodiment when executing the computer program.
[0117] See also Figure 4 The electronic device further includes: a bus 403 and a communication interface 404, a processor 402, a communication interface 404 and a memory 401 are connected via the bus 403; the processor 402 is used to execute an executable module stored in the memory 401, such as a computer program.
[0118] The memory 401 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 404 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0119] The bus 403 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0120] Among them, the memory 401 is used to store programs, and the processor 402 executes the program after receiving the execution instruction. The method executed by the device defined by the process disclosed in any embodiment of the present application can be applied to the processor 402 or implemented by the processor 402.
[0121] The processor 402 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 402. The above processor 402 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined to execute. The software module may be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 401, and the processor 402 reads the information in the memory 401 and completes the steps of the above method in combination with its hardware.
[0122] Corresponding to the above-mentioned vehicle-mounted image rendering method based on AI algorithm, an embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to execute the steps of the above-mentioned vehicle-mounted image rendering method based on the YTS engine AI algorithm.
[0123] The vehicle-mounted image rendering device based on the YTS engine AI algorithm provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The device provided in the embodiment of the present application, its implementation principle and the technical effect produced are the same as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0124] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0125] For another example, the flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the device, method and computer program product according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0126] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0128] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or partly be embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the vehicle-mounted image rendering method based on the YTS engine AI algorithm described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0129] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0130] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solution of the present application, rather than to limit it. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solution recorded in the aforementioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solution deviate from the scope of the technical solution of the embodiment of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A vehicle-mounted image rendering method based on the YTS engine AI algorithm, characterized in that: The method comprises: Obtain vehicle video data; Identify a first subject object in the vehicle-mounted video data by using an AI algorithm and analyze a dynamic action corresponding to the first subject object in the vehicle-mounted video data; Determining a target frame containing the execution process corresponding to the dynamic action from the multiple frames corresponding to the vehicle-mounted video data; For each target frame in the plurality of target frames, determining an action frame position range corresponding to the dynamic action performed by the first subject object in the target frame; Determine the next second target frame of each first target frame according to the sequential frame playback sequence of the target frame in the vehicle video data, and determine the first action picture position range corresponding to the first target frame and the second action picture position range corresponding to the second target frame based on the action picture position range corresponding to each target frame; wherein the first target frame is any target frame in the multiple frames corresponding to the vehicle video data that contains the corresponding execution process of the dynamic action; Under the condition that the relative positions of the first target frame picture and the second target frame picture are kept unchanged, determining a third action picture position range corresponding to the second action picture position range in the first target frame picture, and determining the first action picture position range and the third action picture position range as important focus position ranges in the first target frame picture; When the in-vehicle video data is played to the first target frame, the important focus position range is enhancedly rendered; the screen distance between each position in the first target frame and the important focus position range is determined; the weakened rendering degree of each position is determined according to the screen distance; the longer the screen distance, the greater the weakened rendering degree; according to the weakened rendering degree of each position, the screen position range other than the important focus position range in the first target frame is weakened according to the corresponding weakened rendering degree, so as to guide the in-vehicle user to focus on the important focus position range and the peripheral position range of the important focus position range during the playback of the first target frame.
2. The method according to claim 1, characterized in that After determining the first action picture position range and the third action picture position range as the important focus position range in the first target frame picture, the method further includes: When the in-vehicle video data is played to the first target frame picture, the important focus position range is rendered according to the first specified rendering method, and the picture position range other than the important focus position range in the first target frame picture is rendered according to the second specified rendering method, so as to guide the in-vehicle user to focus on the first action picture position range during the playback of the first target frame picture and to focus on the second action picture position range in advance through the third action picture position range during the playback from the first target frame picture to the second target frame picture; Among them, the first specified rendering method includes at least one of ray tracing rendering, color rendering and animation style rendering; the second specified rendering method includes black and white rendering and a rendering method that keeps the picture style unchanged.
3. The method according to claim 2, characterized in that The rendering of the important focus position range according to a first specified rendering mode, and the rendering of a screen position range other than the important focus position range in the first target frame screen according to a second specified rendering mode, includes: Decomposing the composite rendering task to be rendered into a plurality of sub-rendering tasks, and setting the plurality of sub-rendering tasks in a task queue; wherein the plurality of sub-rendering tasks include a first sub-rendering task for rendering the important focus position range according to a first specified rendering method, and a second sub-rendering task for rendering a screen position range other than the important focus position range in the first target frame screen according to a second specified rendering method; The plurality of sub-rendering tasks are dynamically allocated to a plurality of threads according to the priorities and dependencies of the tasks, and the plurality of sub-rendering tasks are executed by the plurality of threads.
4. The method according to claim 1, characterized in that: After determining the target frame image including the execution process corresponding to the dynamic action from the multiple frames corresponding to the vehicle-mounted video data, the method further includes: Performing image recognition on the target frame to obtain an image recognition result, and judging whether the target first subject object in the target frame conforms to a specified logical rule according to the image recognition result; If the target first main object does not conform to the specified logical rule, the target first main object is eliminated to obtain a target frame picture after the elimination process.
5. The method according to claim 1, characterized in that After acquiring the vehicle-mounted video data, the method further includes: Acquire film and television resource image data from the vehicle-mounted video data, and identify a second subject object in the film and television resource image data; Analyze the film and television resource image data and the style data of the second subject object through the YTS engine; For the second main object in a plurality of different specified scenes, generating object appearance data corresponding to the second main object according to the style data; The second main object is rendered based on the object appearance data to display the second main object in accordance with the style corresponding to the specified scene.
6. The method according to claim 1, characterized in that The first subject object includes any one or more of the following: Vehicle objects, character objects, and static object objects.
7. A vehicle-mounted image rendering device based on the YTS engine AI algorithm, characterized in that: include: An acquisition module, used to acquire vehicle-mounted video data; An identification module, configured to identify a first subject object in the vehicle-mounted video data by using an AI algorithm and analyze a dynamic action corresponding to the first subject object in the vehicle-mounted video data; A first determination module is used to determine a target frame containing the execution process corresponding to the dynamic action from multiple frames corresponding to the vehicle-mounted video data; A second determination module is used to determine, for each target frame in the plurality of target frames, an action picture position range corresponding to the dynamic action performed by the first subject object in the target frame; A third determination module is used to determine the next second target frame of each first target frame according to the sequential frame playback sequence of the target frames played in the vehicle-mounted video data, and determine the first action picture position range corresponding to the first target frame and the second action picture position range corresponding to the second target frame based on the action picture position range corresponding to each target frame; wherein the first target frame is any target frame in the multiple frames corresponding to the vehicle-mounted video data that contains the corresponding execution process of the dynamic action; a fourth determination module, configured to determine a third action picture position range corresponding to the second action picture position range in the first target frame picture while keeping the relative overall picture positions of the first target frame picture and the second target frame picture unchanged, and determine the first action picture position range and the third action picture position range as important focus position ranges in the first target frame picture; A rendering module, used for enhancing the rendering of the important focus position range when the vehicle-mounted video data is played to the first target frame; The rendering module is also used to: determine the picture distance between each position in the first target frame picture and the important focus position range; determine the weakened rendering degree of each position according to the picture distance; the longer the picture distance, the greater the weakened rendering degree; according to the weakened rendering degree of each position, weaken the rendering of the picture position range other than the important focus position range in the first target frame picture according to the corresponding weakened rendering degree, so as to guide the in-vehicle user to focus on the important focus position range and the surrounding position range of the important focus position range during the playback of the first target frame picture.
8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video frame action feature extraction method and device
CN114973410A
Cloud rendering method and device
CN117576358A