Differential quantization method and device of sensing system, electronic equipment and storage medium

By acquiring video images of the camera system in an indoor environment and analyzing the imaging system configuration parameters using a target detection model, the problem of difficulty in isolating the influence of a single variable in the perception performance evaluation of the camera system is solved, enabling accurate perception performance evaluation and rapid iteration.

CN121284201APending Publication Date: 2026-01-06SHANGHAI PHIGENT QIJI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511314633.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing technologies cannot accurately isolate the influence of a single variable in the perception performance evaluation of camera systems, resulting in inaccurate evaluation results and low repeatability, making it difficult to control the influence of environmental variables in real physical environments.

Method used

In an indoor environment, video images captured by two camera systems are acquired, ensuring that the camera parameters are consistent. The images are then processed using a target detection model to analyze the impact of different imaging system configuration parameters on perception performance.

Benefits of technology

It achieves accurate quantitative comparison of sensing distance, confidence level, and category, avoids environmental noise interference, improves the difference quantification accuracy of sensing system, shortens the testing cycle, saves manpower and equipment resources, and is suitable for rapid iteration and indoor environment testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284201A_ABST
    Figure CN121284201A_ABST
Patent Text Reader

Abstract

The invention provides a difference quantification method and device of a sensing system, electronic equipment and a storage medium. Comprising the following steps: acquiring a first video image and a second video image shot by a first camera system and a second camera system respectively; the first camera shooting system and the second camera shooting system are systems with single imaging system configuration variables and have the same camera parameters, and the first video image and the second video image are obtained by the first camera shooting system and the second camera shooting system through shooting simulation scene videos played by the display device in an indoor scene; processing the first video image and the second video image based on a target detection model to obtain a first target detection result of the target object in the first video image and a second target detection result of the target object in the second video image; and according to the first target detection result and the second target detection result, determining an influence result of different imaging system configuration parameters on the perception performance of the perception system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of performance evaluation technology, and in particular to a method, apparatus, electronic device and storage medium for quantifying differences in a sensing system. Background Technology

[0002] The perception performance evaluation of camera systems plays a crucial role in the field of autonomous driving. Currently, the evaluation of camera system perception performance typically involves deploying multiple camera systems (e.g., mounted on vehicles or fixed equipment) in a real physical environment, allowing them to simultaneously or sequentially capture the same scene (e.g., urban roads, nighttime lighting). Then, a pre-trained perception model is used to infer from the captured images or videos, extracting object detection results (including bounding boxes, confidence scores, and categories). These results are then compared manually or automatically to evaluate the performance differences between different camera systems (e.g., different ISP (Image Signal Processor) parameters or hardware chips).

[0003] In actual road testing, it is difficult to strictly control environmental variables (such as lighting, weather, and object movement), which leads to unexpected differences in images captured by different camera systems. As a result, it is impossible to accurately isolate the influence of a single variable (such as ISP software parameters or hardware chips), and the evaluation results are affected by the environment, resulting in inaccuracy and low repeatability. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for quantifying differences in a sensing system, in order to solve the problems that existing technologies cannot accurately isolate the influence of a single variable, and that evaluation results are inaccurate and have low repeatability due to environmental influences.

[0005] To solve the above-mentioned technical problems, the embodiments of this application are implemented as follows:

[0006] In a first aspect, embodiments of this application provide a method for quantifying differences in a sensing system, the method comprising:

[0007] The system acquires a first video image and a second video image captured by a first camera system and a second camera system, respectively. The first camera system and the second camera system are systems with a single imaging system configuration variable. The first camera system and the second camera system have the same camera parameters. The first video image and the second video image are respectively obtained by the first camera system and the second camera system capturing simulated scene videos played on a display device in an indoor scene.

[0008] The first video image and the second video image are processed based on the target detection model to obtain the first target detection result of the target object in the first video image and the second target detection result of the target object in the second video image.

[0009] Based on the first target detection result and the second target detection result, the impact of different imaging system configuration parameters on the perception performance of the perception system is determined.

[0010] Secondly, embodiments of this application provide a differential quantification device for a sensing system, the device comprising:

[0011] The image acquisition module is used to acquire a first video image and a second video image captured by the first camera system and the second camera system, respectively; the first camera system and the second camera system are systems with a single imaging system configuration variable, the first camera system and the second camera system have the same camera parameters, and the first video image and the second video image are respectively: captured by the first camera system and the second camera system in an indoor scene from a simulated scene video played by a display device;

[0012] The result acquisition module is used to process the first video image and the second video image based on the target detection model to obtain the first target detection result of the target object in the first video image and the second target detection result of the target object in the second video image.

[0013] The performance determination module is used to determine the impact of different imaging system configuration parameters on the perception performance of the perception system based on the first target detection result and the second target detection result.

[0014] Thirdly, embodiments of this application provide an electronic device, including:

[0015] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the differential quantization method of the sensing system described in any of the preceding claims.

[0016] Fourthly, embodiments of this application provide a readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the differential quantization method of the sensing system described in any of the preceding claims.

[0017] In this embodiment, by having different camera systems capture the same pre-recorded video, the input environment is ensured to be completely consistent, isolating the influence of a single variable. This enables precise quantitative comparison of perception distance, confidence level, and category, avoiding environmental noise interference during field testing and improving the accuracy of difference quantification of perception systems. No outdoor deployment is required; only indoor video playback and recording are needed, shortening the testing cycle to a few hours. This is suitable for rapid iteration (such as ISP parameter optimization), saving manpower and equipment resources. Furthermore, it is applicable to indoor environments, avoiding the safety risks of field testing. It also supports multiple variable scenario expansions (such as different resolutions or lighting conditions), providing a reliable basis for camera system optimization and promoting advancements in perception technologies in fields such as autonomous driving and surveillance.

[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating the steps of a differential quantification method for a sensing system provided in this application embodiment;

[0021] Figure 2 A flowchart illustrating the steps of a video image acquisition method provided in this application embodiment;

[0022] Figure 3 A flowchart illustrating the steps of another video image acquisition method provided in this application embodiment;

[0023] Figure 4 A flowchart illustrating the steps of a method for determining the impact of perception performance in an embodiment of this application;

[0024] Figure 5 A flowchart illustrating the steps of a method for obtaining influence parameters provided in this application embodiment;

[0025] Figure 6 A flowchart illustrating the steps of another differential quantification method for a sensing system provided in this application embodiment;

[0026] Figure 7This is a schematic diagram of the structure of a difference quantification device for a sensing system according to an embodiment of this application;

[0027] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] Reference Figure 1 The diagram illustrates a flowchart of the steps in a differential quantification method for a sensing system provided in an embodiment of this application. Figure 1 As shown, the differential quantification method of the sensing system may include steps 101, 102 and 103.

[0030] Step 101: Acquire the first video image and the second video image captured by the first camera system and the second camera system respectively; the first camera system and the second camera system are systems with a single imaging system configuration variable, the first camera system and the second camera system have the same camera parameters, and the first video image and the second video image are respectively: captured by the first camera system and the second camera system in an indoor scene from a simulated scene video played by a display device.

[0031] In this embodiment, the first camera system and the second camera system are two independent camera systems used to capture video images. Their core characteristic is that they differ only in a single imaging system configuration variable; all other key camera parameters are identical, allowing for the independent study of the impact of this single configuration variable on perception performance. The imaging system configuration parameter refers to either the image signal processing (ISP) parameters or the hardware image sensor chip parameters.

[0032] In this example, the first camera system and the second camera system can be two camera systems with different ISP parameters but identical other parameters. Alternatively, they can be two camera systems with different hardware image sensor parameters but identical other parameters, such as the two camera systems using X8B and IMX728 image sensors respectively.

[0033] Camera parameters refer to the basic parameters that the first and second camera systems maintain consistency, used to eliminate other interfering factors. In this example, camera parameters may include, but are not limited to, focal length, resolution, etc.

[0034] Indoor scenes refer to video images taken in enclosed or semi-enclosed indoor spaces. Environmental control (such as fixed lighting intensity and a clean background) is required to ensure the stability of shooting conditions and reduce the impact of environmental interference on experimental results.

[0035] Display devices refer to standardized equipment that plays "simulated scene videos" and serves as the subject of a camera system. Common devices include: professional monitors (such as 4K HDR monitors), projectors, LED screens, etc. It is necessary to ensure the consistency of display brightness, color, and refresh rate, and to avoid errors introduced by differences in the devices themselves.

[0036] Simulated scene videos refer to standardized video content played on display devices, containing target objects to be detected (such as pedestrians, vehicles, small objects, moving targets, etc.), used to provide uniform shooting material for camera systems. The videos must possess diversity: diverse target categories (people, vehicles, objects), varied target postures (front / side, stationary / moving), and controllable scene complexity (simple / complex backgrounds) to comprehensively evaluate perception performance.

[0037] In this example, the simulated scene may include at least one of a variety of typical driving scenarios, such as a high color temperature scene (e.g., a sunny midday with high brightness and uniform illumination), a low color temperature scene (e.g., a warm-toned twilight scene), a low light noise scene (e.g., a low-light night), and a scene with dynamic changes in illumination (e.g., a high-dynamic scene of passing through a tunnel with rapid changes in light and darkness).

[0038] In practical applications, the preparation of simulated scene videos can be carried out first, which is the video data preparation stage, and can include: the raw video data acquisition stage and the system setting stage.

[0039] First, prepare a video source, i.e., record or acquire a special high-definition video (resolution 3840*2160, frame rate no less than 30fps). The video content should cover a variety of typical scenes, including a sunny midday scene (high brightness and uniform lighting) - a high color temperature representative scene, a dusky scene (warm-toned light) - a low color temperature representative scene, a nighttime scene (low light and noise) and a high dynamic range scene (rapid changes in light and dark) passing through a tunnel. The video should include various target objects (such as vehicles, pedestrians, and road signs) to simulate a realistic perceived environment. The video length should be no less than 5 minutes to ensure scene diversity and continuity.

[0040] Then, you can prepare two camera systems, ensuring they have only a single variable. For example, system A uses an X8B chip, and system B uses an IMX728 chip. Or, they could use the same hardware chip but different ISP parameters, etc.

[0041] When testing the impact of a single imaging system configuration variable on the perception performance of a sensing system, a first video image and a second video image captured by a first camera system and a second camera system, respectively, can be obtained. The first video image and the second video image can be obtained by the first camera system and the second camera system capturing a simulated scene video played on a display device within an indoor scene.

[0042] In one specific implementation of this application, the first camera system and the second camera system can be placed in the same indoor scene of a simulated driving environment, and the display device can be controlled to play a simulated scene video. Then, the first camera system and the second camera system simultaneously focus on the playing simulated scene video to capture video. This implementation process can be combined with... Figure 2 The following is a detailed description.

[0043] Reference Figure 2 The diagram illustrates a flowchart of a video image acquisition method provided in an embodiment of this application. Figure 2 As shown, the video image acquisition method may include steps 201 and 202.

[0044] Step 201: When the first camera system and the second camera system are placed in the same indoor scene, control the display device to play a pre-recorded video simulating the perceived environment, the video containing the target object.

[0045] In this embodiment, the video simulating the perception environment refers to a pre-recorded video used to simulate the real-world environment faced by the vehicle's perception system during driving. The video content must include typical elements of the driving scenario, such as other vehicles, pedestrians, traffic lights, lane lines, and obstacles, and must have dynamic changing characteristics (such as vehicle movement and pedestrians crossing the road) to simulate the perception scenario in real driving.

[0046] When testing the impact of a single imaging system configuration variable on the perception performance of a sensing system, two systems (i.e., the first camera system and the second camera system) can be placed in the same indoor environment (e.g., a darkroom), and the aforementioned video can be played on a high-definition display screen (e.g., a 4K monitor). Furthermore, the camera parameters (e.g., focal length, resolution) of the two systems should be kept consistent to ensure fairness in the capture process.

[0047] In practice, both systems should ensure that their height and distance from the display device remain consistent.

[0048] The display device (such as a high-resolution professional monitor or LED video wall) can be fixed in a preset position within the scene, ensuring that its display surface faces the shooting direction of the camera system, and that its height and viewing angle conform to the typical observation angle of a vehicle perception system in simulated driving (e.g., 1.2-1.5 meters above the ground and 3-5 meters horizontally from the camera system). Then, pre-recorded simulated perception environment videos can be loaded onto the display device. The videos should cover various driving scene segments (e.g., city roads, highways, intersections), including different types of target objects and dynamic behaviors. The video is controlled by playback software to loop at a fixed frame rate (e.g., 30fps), ensuring that the video is smooth and free of stuttering or screen tearing.

[0049] Step 202: Obtain the first video image and the second video image captured by the first camera system and the second camera system respectively when both are focused on the video displayed on the display device.

[0050] Focusing on the video displayed on the display device refers to the state in which the camera system, by adjusting parameters such as lens focal length, ensures that the simulated perceived environment video played on the display device is clearly imaged on the camera system's imaging sensor. This ensures that target objects (such as vehicles and pedestrians) in the video have clear edges and discernible details in the captured video image, providing high-quality image data for subsequent target detection.

[0051] Both a first camera system and a second camera system can be used to focus on the video displayed on the display device and capture video images separately, thereby obtaining a first video image and a second video image. That is, provided that the first and second camera systems are accurately focused, video images captured by the two camera systems are acquired simultaneously, ensuring that only a single imaging system configuration variable differs.

[0052] In practical applications, the first and second camera systems can be fixed to preset positions in the simulated driving environment's indoor scene using a tripod. Ensure that the installation height and horizontal angle of both cameras are completely consistent, that the shooting direction is directly facing the display surface of the display device, and that the center of the lens is on the same horizontal line as the center of the display device, guaranteeing a field-of-view overlap of ≥95%. Focusing operations are then performed on both camera systems. Using the cameras' built-in autofocus function or manually adjusting the focus, the simulated perception environment video played on the display device is clearly presented in the real-time preview, with a focus on ensuring that the outlines and details of target objects (such as vehicles and pedestrians) in the image are clearly distinguishable. The focusing effect can be verified by capturing test frames and magnifying them to check the sharpness of the target edges.

[0053] Once focusing is complete and the display device is playing the video normally, the image acquisition functions of both camera systems are started simultaneously, and the shooting timestamp is recorded to ensure the time synchronization of the two sets of video images. During the acquisition process, the device status is continuously monitored to ensure that there are no interruptions in recording or frame drops. The acquired first and second video images are stored in a standardized format (such as MP4) and labeled with the corresponding camera system identifier.

[0054] This application provides precise and controllable experimental data support for analyzing the impact of imaging configuration on perception performance by constructing a standardized simulated driving environment and controlling the acquisition of images by a single variable, thereby improving the scientificity and reliability of perception performance evaluation.

[0055] In another specific implementation of this application, virtual driving environment information can be pre-created and input into two sets of camera systems. These two camera systems then capture video in an indoor scene to obtain a first video image and a second video image. This implementation process can be combined with... Figure 3 The following is a detailed description.

[0056] Reference Figure 3 The diagram illustrates a flowchart of another video image acquisition method provided in an embodiment of this application. Figure 3 As shown, the video image acquisition method may include steps 301, 302 and 303.

[0057] Step 301: Obtain pre-created virtual driving environment information, and render and input the virtual driving environment information into the first camera system and the second camera system.

[0058] In this embodiment, virtual driving environment information refers to a set of digital environmental data pre-created using computer graphics technology to simulate real driving scenarios. It includes environmental parameters of the virtual scene (light intensity, weather conditions, time dimension, etc.) and scene physical rules (such as light reflection patterns, etc.), which can be converted into visual images through rendering technology and input to the camera system.

[0059] In practical applications, computer simulation software can be used to generate virtual scenes: for example, using Unity or CARLA simulation engines to create virtual environments (including scenes such as sunny days and dusk), and then rendering them into a camera system.

[0060] When testing the impact of a single imaging system configuration variable on the perception performance of the perception system, pre-created virtual driving environment information can be obtained, and the virtual driving environment information can be rendered and input into the first camera system and the second camera system.

[0061] Step 302: When the first camera system and the second camera system are placed in the same indoor scene, control the display device to play a pre-recorded video of the simulated perception environment, the video containing the target object.

[0062] After the virtual driving environment information is rendered and input into the first camera system and the second camera system, the first camera system and the second camera system can be placed in the same indoor scene, and the display device can be controlled to play a pre-recorded video of the simulated perception environment, and the video playing contains the target object.

[0063] Step 303: Obtain the first video image and the second video image captured by the first camera system and the second camera system respectively when both are focused on the video displayed on the display device.

[0064] Once focusing is complete and the display device is playing the video normally, the image acquisition functions of both camera systems are started simultaneously, and the shooting timestamp is recorded to ensure the time synchronization of the two sets of video images. During the acquisition process, the device status is continuously monitored to ensure that there are no interruptions in recording or frame drops. The acquired first and second video images are stored in a standardized format (such as MP4) and labeled with the corresponding camera system identifier.

[0065] This application embodiment uses a combination of virtual and physical scene input methods to collect image data from differentiated camera systems under controllable conditions. This provides multi-dimensional and highly consistent experimental evidence for accurately analyzing the impact of imaging parameters on perception performance, thereby improving the comprehensiveness and accuracy of perception system evaluation.

[0066] After acquiring the first video image and the second video image captured by the first camera system and the second camera system respectively, step 102 is executed.

[0067] Step 102: Process the first video image and the second video image based on the target detection model to obtain the first target detection result of the target object in the first video image and the second target detection result of the target object in the second video image.

[0068] An object detection model refers to an object recognition algorithm model based on deep learning or traditional algorithms, used to locate and identify target objects from video images. The input is a video frame image, and the output is key information about the target (such as detection box, category, and confidence score).

[0069] The first and second object detection results refer to the outputs of the object detection model after processing the first and second video images, respectively. In this example, both object detection results include: object localization information (i.e., the bounding box of the object in the image), object category (i.e., the detected object belongs to a preset category (such as "pedestrian", "car", etc.)), and confidence score (the model's confidence score for the detection result (between 0 and 1)).

[0070] After obtaining the first video image and the second video image, the first video image and the second video image can be processed based on the target detection model to obtain the first target detection result of the target object in the first video image and the second target detection result of the target object in the second video image.

[0071] In practical applications, a pre-trained YOLOv5n object detection model (a lightweight version suitable for real-time inference) can be used, combined with the OpenVINO framework for optimization and acceleration. The model has been trained on datasets or similar datasets and supports common categories (such as people, vehicles, and animals).

[0072] The inference process can be as follows: input two recorded video clips into the model for inference. For each frame of the image, the model outputs a perception result, including:

[0073] Bounding Box: The position and size of the target.

[0074] Confidence Score: The reliability of the detection (range 0-1).

[0075] Class: Target type (e.g., "vehicle" or "pedestrian").

[0076] Inference uses OpenVINO's inference engine, optimized for operation on Intel hardware, ensuring low latency (<10ms / frame).

[0077] The result of perceptual inference can be a vector of [xmin, ymin, xmax, ymax, confidence, class_id], representing the x-coordinate (pixel value) of the top left corner of the detection box, the y-coordinate (pixel value) of the top left corner of the detection box, the x-coordinate (pixel value) of the bottom right corner of the detection box, the y-coordinate (pixel value) of the bottom right corner of the detection box, the detection confidence (0~1), the larger the value, the more confident the model is, and the class ID (integer), which corresponds to the class index in data / coco.yaml.

[0078] Step 103: Based on the first target detection result and the second target detection result, determine the impact of different imaging system configuration parameters on the perception performance of the perception system.

[0079] The perception performance impact result refers to the conclusions drawn from analyzing the specific impact of a "single imaging system configuration variable" on the perception system performance by comparing the detection results of the first and second targets. Examples include the impact of different ISP parameters on the perception system's performance, or the impact of different hardware photosensitive chips on the perception system's performance. This perception performance impact result can be used to guide hardware / software optimization, possessing significant practical value and economic benefits.

[0080] After obtaining the first and second target detection results, the impact of different imaging system configuration parameters on the perception performance of the sensing system can be determined based on these results. In specific implementations, the impact parameters can be calculated based on the target detection results of the two systems, and the impact on perception performance can be determined based on these calculated parameters. This implementation process can be combined with... Figure 4 The following is a detailed description.

[0081] Reference Figure 4 The diagram illustrates a flowchart of a method for determining the impact of perceived performance according to an embodiment of this application. Figure 4 As shown, the method for determining the impact of perception performance may include steps 401 and 402.

[0082] Step 401: Based on the first detection box, first confidence level, and first category corresponding to each frame of the first video image, and the second detection box, second confidence level, and second category corresponding to each frame of the second video image, determine the influence parameters of different imaging system configuration parameters on the perception system in a single video image pair.

[0083] In this embodiment, the first target detection result may include: the first detection box, the first confidence score, and the first category of the target object corresponding to each frame of the first video image.

[0084] The second target detection result may include: the second detection box, the second confidence score, and the second category of the target object corresponding to each frame of the second video image.

[0085] After obtaining the first and second target detection results, the influence parameters of different imaging system configurations on the perception system in a single video image pair can be determined based on the first detection box, first confidence score, and first category corresponding to each frame of the first video image, and the second detection box, second confidence score, and second category corresponding to each frame of the second video image. That is, based on the first and second target detection results of a single video image pair, the impact of different imaging configurations on the perception system in a single frame is quantified through comparative analysis. The process for obtaining the influence parameters can be combined with... Figure 5 The following is a detailed description.

[0086] Reference Figure 5 The diagram illustrates a flowchart of a method for obtaining influence parameters according to an embodiment of this application. Figure 5 As shown, the method for obtaining the influencing parameters may include steps 501, 502, 503, and 504.

[0087] Step 501: Determine the sensing distance parameter based on the first detection box and the second detection box matched by the first detection box.

[0088] In this embodiment, the perception distance parameter can be determined based on the first detection box and the second detection box that matches the first detection box. It can be understood that, for the first video image and the second video image, a matching single frame of the first video image and the second video image (i.e., video images with the same timestamp) can be determined based on the timestamps of the first video image and the second video image.

[0089] In the specific implementation, the perceptual reasoning result is [xmin, ymin, xmax, ymax, confidence, class_id].

[0090] The linear dimensions of the target object in the image (e.g., height in pixels h) are related to the true distance D as follows:

[0091] Formula: h = f × H / Dh: The imaging height of the object in the image (in pixels). For example, the height of the bounding box extracted from the YOLO target bounding box.

[0092] H: The actual physical height of the object (in meters).

[0093] f: The camera's focal length (in pixels). This can be obtained through the camera's intrinsic matrix, usually derived from camera calibration. For example, for a 1920×1080 resolution camera, f can range from a few hundred to a few thousand pixels (depending on the lens).

[0094] D: The actual distance (in meters) from the object to the camera.

[0095] The sensing distance parameter is calculated as follows:

[0096] Calculate h = ymax - ymin (pixels).

[0097] D is calculated using the known H and f.

[0098] Compare the first camera system and the second camera system, and calculate the difference (e.g., ΔD = |D_A - D_B|).

[0099] Quantify the impact of variables on the detection of near and far targets (e.g., system A detects targets at a distance 20% farther than system B).

[0100] Step 502: Determine the confidence difference parameter based on the first confidence level and the second confidence level matched by the first confidence level.

[0101] The confidence difference parameter can be determined using a first confidence level and a matched second confidence level. Specifically, confidence level: the difference in confidence levels for the same objective (e.g., an average difference > 0.1 indicates a significant effect).

[0102] Step 503: Determine the category matching rate based on the first category and the second category that matches the first category.

[0103] Simultaneously, the category matching rate can be determined based on the first category and the matched second category. Specifically, the category matching rate is statistically analyzed (e.g., misclassification rate <5%). In practical applications, the matching rate can be calculated through semantic similarity.

[0104] Step 504: Use the perception distance parameter, the confidence difference parameter, and the category matching rate as parameters to determine the influence of different imaging system configuration parameters on the perception system in a single video image pair.

[0105] After obtaining the perceptual distance parameter, confidence difference parameter, and class matching rate of a single video image pair, these parameters can be used as parameters influencing the perception system on the different imaging system configuration parameters within the single video image pair.

[0106] This application embodiment calculates results by statistical comparison of multi-dimensional indicators, classifies results by scenario, supports visualization, and can improve the quantitative accuracy of configuration parameters of different imaging systems.

[0107] Step 402: Based on the influence parameters of multiple video image pairs, determine the impact of different imaging system configuration parameters on the perception performance of the perception system.

[0108] After obtaining the influence parameters of multiple video image pairs, the impact of different imaging system configuration parameters on the perception performance of the perception system can be determined based on these parameters. Specifically, multi-frame results can be aggregated, statistical indicators (such as mean difference and standard deviation) can be calculated, and reports can be output according to scene classification (sunny day, dusk, etc.). Visualization tools (such as Matplotlib) can be used to generate heatmaps or curves to show the specific impact of variables on perception (e.g., "IMX728 chip improves confidence by 15% in low-light scenes").

[0109] The embodiments of this application avoid the subjectivity of manual evaluation by standardizing the calculation of influence parameters and multi-frame statistical analysis, making the evaluation of the impact of imaging configuration on perception performance more objective and reliable.

[0110] In this embodiment, video images from two camera systems can be generated using a generative adversarial network (GAN), followed by a differential quantization process. This implementation can be combined with... Figure 6 The following is a detailed description.

[0111] Reference Figure 6 This illustrates a flowchart of another differential quantification method for a sensing system provided in an embodiment of this application. Figure 6 As shown, the differential quantification method of the sensing system may include steps 601, 602, 603 and 604.

[0112] Step 601: Obtain the configuration parameters of the first imaging system and the parameters of the first camera corresponding to the first camera system, and the configuration parameters of the second imaging system and the parameters of the second camera corresponding to the second camera system.

[0113] In this embodiment, the configuration parameters of the first imaging system and the parameters of the first camera corresponding to the first camera system, as well as the configuration parameters of the second imaging system and the parameters of the second camera corresponding to the second camera system, can be obtained.

[0114] The configuration parameters of the first imaging system and the second imaging system are different, while the parameters of the first camera and the second camera are the same.

[0115] Step 602: Based on the generative adversarial network, process the configuration parameters of the first imaging system, the parameters of the first camera and the target driving environment, and the configuration parameters of the second imaging system, the parameters of the second camera and the target driving environment to obtain the third video image corresponding to the first camera system and the fourth video image corresponding to the second camera system.

[0116] Target driving environment parameters refer to the environmental characteristics used to describe the driving scenario, such as light intensity (strong light / weak light / backlight), weather conditions (sunny / rainy / foggy / snowy), scene type (city road / highway / tunnel), dynamic interference (glare / shadow), etc.

[0117] Generative Adversarial Networks (GANs) are deep learning models used to simulate imaging effects. They consist of a generator (which generates simulated images based on input parameters) and a discriminator (which distinguishes between real and generated images). Through adversarial training, they generate video images that are highly similar to real scenes.

[0118] After obtaining the imaging system configuration parameters and camera parameters of the two camera systems, a generative adversarial network (GAN) can be used to process the configuration parameters of the first imaging system, the parameters of the first camera, and the target driving environment parameters, as well as the configuration parameters of the second imaging system, the parameters of the second camera, and the target driving environment parameters, respectively, to obtain the third video image corresponding to the first camera system and the fourth video image corresponding to the second camera system. Specifically, the parameters of the first camera system and the target driving environment parameters are input into the GAN generator to generate the third video image; similarly, the parameters of the second camera system and the same environmental parameters are used to generate the fourth video image (ensuring that only the imaging system configuration parameters differ, while the environment remains consistent).

[0119] Step 603: Process the third video image and the fourth video image based on the target detection model to obtain the third target detection result of the target object in the third video image and the fourth target detection result of the target object in the fourth video image.

[0120] After obtaining the third and fourth video images, they can be processed based on the object detection model to obtain the third object detection result for the target object in the third video image and the fourth object detection result for the target object in the fourth video image. That is, the third and fourth video images are input into the object detection model respectively to obtain the corresponding third and fourth object detection results (including the target's bounding box, confidence score, and category).

[0121] Step 604: Based on the third target detection result and the fourth target detection result, determine the impact of different imaging system configuration parameters on the perception performance of the perception system.

[0122] After obtaining the third and fourth target detection results, the impact of different imaging system configuration parameters on the perception performance of the perception system can be determined based on these results.

[0123] This application's embodiments generate simulated images using GANs, replacing some real-vehicle road tests and reducing the cost and time required for hardware deployment and scenario building. Simultaneously, it allows for precise control of target driving environment parameters, eliminating interference from environmental differences and enabling a purer comparison of the impact of different camera system parameters. It can quickly evaluate the effectiveness of different imaging system configuration parameters, providing data support for optimal parameter selection and improving the robustness of the perception system in complex driving environments.

[0124] The differential quantification method for perception systems provided in this application ensures a completely consistent input environment by having different camera systems capture the same pre-recorded video. This isolates the influence of a single variable and enables precise quantitative comparison of perception distance, confidence level, and category. It avoids environmental noise interference during field testing, thus improving the differential quantification accuracy of the perception system. No outdoor deployment is required; only indoor video playback and recording are needed. The testing cycle is shortened to a few hours, making it suitable for rapid iteration (such as ISP parameter optimization) and saving manpower and equipment resources. Furthermore, it is applicable to indoor environments, avoiding the safety risks of field testing. It also supports various variable scenario expansions (such as different resolutions or lighting conditions), providing a reliable basis for camera system optimization and promoting advancements in perception technologies in fields such as autonomous driving and surveillance.

[0125] Reference Figure 7 The diagram shows a structural schematic of a difference quantification device for a sensing system provided in an embodiment of this application. Figure 7 As shown, the difference quantization device 700 of the sensing system may include the following modules:

[0126] The image acquisition module 710 is used to acquire a first video image and a second video image captured by the first camera system and the second camera system, respectively; the first camera system and the second camera system are systems with a single imaging system configuration variable, the first camera system and the second camera system have the same camera parameters, and the first video image and the second video image are respectively: captured by the first camera system and the second camera system in an indoor scene from a simulated scene video played by a display device;

[0127] The result acquisition module 720 is used to process the first video image and the second video image based on the target detection model to obtain a first target detection result of the target object in the first video image and a second target detection result of the target object in the second video image.

[0128] The performance determination module 730 is used to determine the impact of different imaging system configuration parameters on the perception performance of the perception system based on the first target detection result and the second target detection result.

[0129] Optionally, the image acquisition module includes:

[0130] The video playback unit is used to control the display device to play a pre-recorded video simulating a perceived environment when the first camera system and the second camera system are placed in the same indoor scene, wherein the target object is included in the frame of the video.

[0131] An image capturing unit is used to acquire the first video image and the second video image captured by the first camera system and the second camera system respectively when both are focused on the video displayed on the display device.

[0132] Optionally, the image acquisition module includes:

[0133] A virtual environment input unit is used to acquire pre-created virtual driving environment information and render and input the virtual driving environment information to the first camera system and the second camera system;

[0134] The video display unit is used to control the display device to play a pre-recorded video simulating a perceived environment when the first camera system and the second camera system are placed in the same indoor scene, wherein the target object is included in the frame of the video.

[0135] The video capturing unit is used to acquire the first video image and the second video image captured by the first camera system and the second camera system respectively when both are focused on the video displayed on the display device.

[0136] Optionally, the first target detection result includes: a first detection box, a first confidence score, and a first category of the target object corresponding to each frame of the first video image;

[0137] The second target detection result includes: the second detection box, the second confidence level, and the second category of the target object corresponding to each frame of the second video image.

[0138] Optionally, the performance determination module includes:

[0139] The influence parameter determination unit is used to determine the influence parameters of different imaging system configuration parameters on the perception system in a single video image pair based on the first detection box, first confidence level and first category corresponding to each frame of the first video image, and the second detection box, second confidence level and second category corresponding to each frame of the second video image.

[0140] The influence result determination unit is used to determine the influence result of different imaging system configuration parameters on the perception performance of the perception system based on the influence parameters of multiple video image pairs.

[0141] Optionally, the influencing parameter determination unit includes:

[0142] The sensing distance determination subunit is used to determine the sensing distance parameters based on the first detection box and the second detection box matched by the first detection box;

[0143] The difference parameter determination subunit is used to determine the confidence difference parameter based on the first confidence level and the second confidence level matched by the first confidence level;

[0144] The matching rate determination subunit is used to determine the category matching rate based on the first category and the second category that the first category matches;

[0145] The influence parameter acquisition subunit is used to take the perception distance parameter, the confidence difference parameter and the category matching rate as the influence parameters of different imaging system configuration parameters on the perception system in a single video image pair.

[0146] Optionally, the imaging system configuration parameters include either image signal processor parameters or hardware photosensitive chip parameters;

[0147] The simulated scenarios include at least one of the following: high color temperature scenario, low color temperature scenario, weak light noise scenario, and dynamic lighting change scenario.

[0148] Optionally, the device further includes:

[0149] The parameter acquisition module is used to acquire the first imaging system configuration parameters and the first camera parameters corresponding to the first camera system, and the second imaging system configuration parameters and the second camera parameters corresponding to the second camera system.

[0150] The video image acquisition module is used to process the configuration parameters of the first imaging system, the parameters of the first camera and the parameters of the target driving environment, and the configuration parameters of the second imaging system, the parameters of the second camera and the parameters of the target driving environment based on the generative adversarial network, respectively, to obtain a third video image corresponding to the first camera system and a fourth video image corresponding to the second camera system.

[0151] The detection result acquisition module is used to process the third video image and the fourth video image based on the target detection model to obtain the third target detection result of the target object in the third video image and the fourth target detection result of the target object in the fourth video image;

[0152] The perception impact result determination module is used to determine the impact of different imaging system configuration parameters on the perception performance of the perception system based on the third target detection result and the fourth target detection result.

[0153] The differential quantification device for the perception system provided in this application ensures a completely consistent input environment by having different camera systems capture the same pre-recorded video. This isolates the influence of a single variable and enables precise quantitative comparison of perception distance, confidence level, and category. It avoids environmental noise interference during field testing, thus improving the differential quantification accuracy of the perception system. No outdoor deployment is required; only indoor video playback and recording are needed. The testing cycle is shortened to a few hours, making it suitable for rapid iteration (such as ISP parameter optimization) and saving manpower and equipment resources. Furthermore, it is applicable to indoor environments, avoiding the safety risks of field testing. It also supports various variable scenario expansions (such as different resolutions or lighting conditions), providing a reliable basis for camera system optimization and promoting advancements in perception technologies in fields such as autonomous driving and surveillance.

[0154] This application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the difference quantization method of the above-described sensing system.

[0155] Figure 8 A schematic diagram of the structure of an electronic device 800 according to an embodiment of the present invention is shown. Figure 8 As shown, the electronic device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 808 into random access memory (RAM) 803. The RAM 803 can also store various programs and data required for the operation of the electronic device 800. The CPU 801, ROM 802, and RAM 803 are interconnected via bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0156] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, microphone, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0157] The various processes and handling described above can be executed by processing unit 801. For example, the methods of any of the above embodiments can be implemented as computer software programs tangibly contained in a computer-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more actions of the methods described above can be performed.

[0158] This application provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the differential quantification method embodiment of the aforementioned sensing system, achieving the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may include, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0159] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0161] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0162] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0163] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0164] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or groups may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0165] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0166] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0167] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0168] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of differential quantization of a perception system, the method comprising: The method comprises: obtaining a first video image and a second video image respectively photographed by a first camera system and a second camera system; the first camera system and the second camera system are systems with a single imaging system configuration variable, the first camera system and the second camera system have the same camera parameters, and the first video image and the second video image are respectively obtained by photographing a simulated scene video played by a display device in an indoor scene by the first camera system and the second camera system; processing the first video image and the second video image based on a target detection model to obtain a first target detection result of a target object in the first video image and a second target detection result of the target object in the second video image; determining an influence result of different imaging system configuration parameters on the perception performance of the perception system according to the first target detection result and the second target detection result.

2. The method of claim 1, wherein, The method comprises: controlling a display device to play a pre-recorded video of a simulated perception environment under the condition that the first camera system and the second camera system are placed in the same indoor scene, and the video contains the target object in the picture; obtaining the first video image and the second video image respectively photographed by the first camera system and the second camera system when they are both focused on the video displayed by the display device.

3. The method of claim 1, wherein, The method comprises: obtaining pre-created virtual driving environment information and rendering the virtual driving environment information into the first camera system and the second camera system; controlling a display device to play a pre-recorded video of a simulated perception environment under the condition that the first camera system and the second camera system are placed in the same indoor scene, and the video contains the target object in the picture; obtaining the first video image and the second video image respectively photographed by the first camera system and the second camera system when they are both focused on the video displayed by the display device.

4. The method of claim 1, wherein, The first target detection result comprises a first detection box, a first confidence and a first category of the target object corresponding to each frame of the first video image; The second target detection result comprises a second detection box, a second confidence and a second category of the target object corresponding to each frame of the second video image.

5. The method of claim 4, wherein, The method comprises: determining an influence parameter of different imaging system configuration parameters on the perception system according to the first detection box, the first confidence and the first category corresponding to each frame of the first video image and the second detection box, the second confidence and the second category corresponding to each frame of the second video image; determining an influence result of different imaging system configuration parameters on the perception performance of the perception system according to the influence parameters of multiple video image pairs.

6. The method of claim 5, wherein, The determining, according to the first detection frame corresponding to each frame of the first video image, the first confidence and the first category, and the second detection frame corresponding to each frame of the second video image, the second confidence and the second category, of an influence parameter of different imaging system configuration parameters on the perception system in a single video image pair, comprises: determining a perception distance parameter according to the first detection frame and the second detection frame matched with the first detection frame; determining a confidence difference parameter according to the first confidence and the second confidence matched with the first confidence; determining a category matching rate according to the first category and the second category matched with the first category; taking the perception distance parameter, the confidence difference parameter and the category matching rate as the influence parameter of different imaging system configuration parameters on the perception system in a single video image pair.

7. The method of claim 1, wherein, The imaging system configuration parameter comprises any one of an image signal processor parameter and a hardware photosensitive chip parameter. The simulation scene comprises at least one of a high color temperature scene, a low color temperature scene, a weak light noise scene and a light dynamic change scene.

8. The method of claim 1, wherein, The method further comprises: obtaining first imaging system configuration parameters and first camera parameters corresponding to a first camera system, and second imaging system configuration parameters and second camera parameters corresponding to a second camera system; processing, based on a generative adversarial network, the first imaging system configuration parameters, the first camera parameters and target driving environment parameters, and the second imaging system configuration parameters, the second camera parameters and the target driving environment parameters respectively, to obtain a third video image corresponding to the first camera system, and a fourth video image corresponding to the second camera system; processing, based on a target detection model, the third video image and the fourth video image to obtain a third target detection result of a target object in the third video image, and a fourth target detection result of the target object in the fourth video image; determining an influence result of different imaging system configuration parameters on the perception performance of the perception system according to the third target detection result and the fourth target detection result.

9. A differential quantization apparatus of a perception system, characterized by, The device comprises: an image acquisition module configured to acquire a first video image and a second video image respectively captured by a first camera system and a second camera system; the first camera system and the second camera system are systems having a single imaging system configuration variable, and the first camera system and the second camera system have the same camera parameters; the first video image and the second video image are respectively obtained by capturing a simulation scene video played by a display device in an indoor scene by the first camera system and the second camera system; a result acquisition module configured to process, based on a target detection model, the first video image and the second video image to obtain a first target detection result of a target object in the first video image, and a second target detection result of the target object in the second video image. A performance determination module is configured to determine an influence result of different imaging system configuration parameters on the perception performance of the perception system according to the first target detection result and the second target detection result.

10. An electronic device, comprising: The application further provides a computer readable storage medium having stored therein instructions of the difference quantification method of the perception system according to any one of claims 1 to 8. The application further provides a computer readable storage medium having stored therein instructions of the difference quantification method of the perception system according to any one of claims 1 to 8.

11. A readable storage medium, characterized by, The application further provides a computer readable storage medium having stored therein instructions of the difference quantification method of the perception system according to any one of claims 1 to 8.