Image intelligent preprocessing and privacy desensitization method and system for VLM reasoning

CN122799072APending Publication Date: 2026-09-22SHANGHAI HENGGE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611273215.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0006]本发明提供一种面向VLM推理的图像智能预处理与隐私脱敏方法及系统,旨在解决现有技术中图像预处理兼容性差、传输效率低、隐私泄露风险高、容错能力不足以及缺乏响应质量闭环控制等问题

Benefits of technology

1.提升图像处理兼容性与传输效率:通过对请求体进行递归遍历精准提取图像数据,并执行等比不放大缩放与格式转换/重新编码,统一了图像规格,避免了高分辨率或非法格式导致的请求被拒或兼容性故障,同时大幅减少了图像体积,降低了网络带宽与Token配额消耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799072A_ABST
    Figure CN122799072A_ABST
Patent Text Reader

Abstract

The application discloses an image intelligent preprocessing and privacy desensitization method and system for VLM reasoning, and relates to the field of cross between software testing and artificial intelligence. The method receives a VLM reasoning request, recursively traverses a request body to identify Base64 image data, performs non-blocking verification of a format whitelist and three-level security restrictions, performs equi-proportional non-scaling and format conversion on the image, performs area mask desensitization processing in response to an assertion type request, records associated metadata and sends the request, performs hallucination detection by cross verification of a named entity recognition model and visual positioning results, generates a response quality score based on hallucination detection, semantic matching and security detection results, collects user feedback, determines samples to be optimized according to image perceptual hashing or processing parameters, and dynamically adjusts verification, preprocessing and desensitization strategies. The method effectively improves image processing compatibility and transmission efficiency, guarantees data privacy and security, and realizes closed-loop optimization of response quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the intersection of software testing engineering and artificial intelligence, and in particular to an image intelligent preprocessing and privacy desensitization method and system for VLM inference. Background Technology

[0002] Visual Language Models (VLMs) are multimodal AI models capable of processing both text and images simultaneously, demonstrating excellent performance in tasks such as image understanding and visual question answering. However, VLMs' powerful semantic analysis capabilities also bring unprecedented privacy risks. Attackers can leverage them to automatically extract private attribute information on a large scale from seemingly harmless images, and protected health information in sensitive fields such as medicine is also at risk of leakage. In AI-driven automated testing platforms, the test execution engine typically calls VLM services to identify interface elements and make assertions, requiring a screenshot of the interface as core input data for each call.

[0003] However, existing technologies have several shortcomings in this image processing and transmission process. Regarding image input specifications, the test devices captured screenshots with extremely wide resolutions. High-resolution screenshots easily exceed the input limits of the VLM service, leading to request rejection or decreased inference quality. Furthermore, the screenshot formats are numerous and inconsistent, with some services only supporting specific formats, easily causing compatibility issues. In terms of transmission efficiency and resource consumption, uncompressed raw screenshots are enormous, consuming significant network bandwidth and token quotas, significantly increasing inference costs. More seriously, assertion evidence screenshots often contain sensitive data such as account information and session identifiers. Combined with VLM's powerful image parsing capabilities, directly sending such data to third-party services poses a high risk of privacy breaches and fails to meet data security compliance requirements.

[0004] Furthermore, existing preprocessing architectures also have significant shortcomings in terms of fault tolerance and decoupling. Existing solutions typically throw exceptions directly and block the entire test execution process when image preprocessing fails, greatly reducing testing efficiency. Simultaneously, the system architecture suffers from field path coupling issues, requiring the specific paths of image fields in the request body to be known and hard-coded in advance. When the request body structure undergoes iterative changes, in-depth modifications to the underlying logic are necessary, resulting in high maintenance costs and poor scalability.

[0005] In summary, this invention proposes an intelligent image preprocessing and privacy desensitization method and system for VLM inference. Summary of the Invention

[0006] This invention provides an intelligent image preprocessing and privacy desensitization method and system for VLM inference, aiming to solve problems such as poor image preprocessing compatibility, low transmission efficiency, high risk of privacy leakage, insufficient fault tolerance, and lack of closed-loop control for response quality in the prior art.

[0007] Firstly, an intelligent image preprocessing and privacy desensitization method for VLM inference includes the following steps: Receive VLM inference requests; The request body of the VLM inference request is recursively traversed, and the strings obtained by the traversal are matched with data URLs to identify the Base64 encoded image data. The identified image data is validated using a format whitelist and security restrictions. Image data that passes the validation is used as images to be preprocessed. If the validation fails or termination is requested, the corresponding skip status and skip reason metadata are recorded in advance. Using the long side of the image as a reference, the image to be preprocessed is subjected to proportional scaling without enlargement and format conversion or re-encoding to obtain a preprocessed image; In response to the identification that the VLM inference request is an assertion request, region masking processing is performed on the preprocessed image based on the masking region to obtain a desensitized image, and the VLM inference request is updated using the desensitized image. Send the updated VLM inference request to the upstream VLM service and obtain a VLM response; Record the original size, processed size, number of bytes, processing parameters, verification status, skipping reason, and associated metadata of the VLM response corresponding to each image; The response entity is extracted from the VLM response, and the response entity is cross-validated with the visual localization results in the original image to obtain the hallucination detection result; based on the hallucination detection result, the semantic matching result between the VLM response and the user request, and the security detection result of the VLM response, a response quality score is generated. Collect user feedback data associated with the VLM response; Based on the associated metadata, the hallucination detection results, the response quality score, and the user feedback data, samples to be optimized are determined according to the dimensions of the same image perceptual hash or the same image processing parameters. Based on the samples to be optimized, at least one of the following is adjusted: the verification strategy of the format whitelist verification and security restriction verification, the preprocessing parameters of the proportional scaling without enlargement and format conversion or re-encoding, and the desensitization strategy of the region masking processing. The adjusted processing strategy is associated with the corresponding image processing features so that subsequent VLM inference requests with the corresponding image processing features adopt the optimized intelligent image preprocessing and privacy desensitization strategy.

[0008] Preferably, the recursive traversal includes: performing a depth-first or breadth-first traversal on the multi-level nested request body based on the data structure type to extract image data; The format whitelist verification includes: identifying the current image format by parsing the file header features or multimedia type identifier of the image data, matching the current image format with a preset legal format feature library, and only allowing successfully matched image data to enter the security restriction verification.

[0009] By adopting the above technical solution, it is possible to adapt to nested structures of any depth in the request body, ensuring that no image data is missed, and to prevent format compatibility failures by blocking illegal formats through MIME type whitelist.

[0010] Preferably, the security restriction verification includes: sequentially executing the image quantity limit, the byte limit after decoding a single image, and the total pixel limit of a single image in a single request; For image data that fails the security limit verification, it is marked as skipped and the skipping reason is recorded, such as exceeding the limit for the number of images, the limit for the number of decoded bytes, the limit for the total number of pixels, or the processing failure.

[0011] By adopting the above technical solution, the three-level security restrictions filter abnormal images layer by layer, and all of them are non-blocking. Images that exceed the restrictions are skipped instead of being directly reported as errors, thus ensuring the continuity of the testing process.

[0012] Preferably, the aspect ratio scaling without enlargement includes: scaling while maintaining the aspect ratio when the longer side of the original image exceeds the maximum longer side; the format conversion or re-encoding includes: generating a preprocessing result according to the quality parameters or compression level parameters corresponding to the target format; the preprocessing parameters include at least one of the maximum longer side, target format, quality parameters, and compression level parameters.

[0013] By adopting the above technical solutions, the proportional non-scaling strategy ensures that the image resolution does not exceed the processing limit of the VLM service, while avoiding unnecessary image quality loss; the unified format conversion improves compatibility.

[0014] Preferably, identifying the VLM inference request as an assertion request includes: identifying it based on the task step type carried in the request, the assertion intent field in the request context, or the call intent obtained based on the request body analysis; the region masking process includes specifying region masking or bottom ratio masking; wherein, for masking coordinate values ​​located in the closed interval of 0 to 1, they are used as normalized ratio values ​​and converted into pixel coordinates; masking coordinate values ​​greater than 1 are used as pixel coordinates, and the resulting masking region is clipped at the boundary.

[0015] By adopting the above technical solutions, the flexible assertion recognition mechanism and dual-mode masking strategy take into account both the coverage and accuracy of automated desensitization, and the dual support of normalized coordinates and pixel coordinates adapts to the needs of different calling scenarios.

[0016] Preferably, when the upstream VLM service returns an HTTP status code greater than or equal to 400 for the updated VLM inference request and is configured to allow retry on failure, the VLM inference request is reconstructed using the original image, the proportional scaling without zooming and the format conversion or re-encoding processing are removed, and when the VLM inference request is an assertion-type request, the region masking processing of the original image is retained. The reconstructed VLM inference request is sent to the upstream VLM service.

[0017] Preferably, the process of extracting response entities from the VLM response and cross-validating the response entities with the visual localization results in the original image to obtain hallucination detection results includes: extracting at least one entity attribute from the VLM response, including object category, quantity, color, and spatial location, using a named entity recognition model; constructing a detection task from the entity attributes and inputting it into the visual localization model to obtain entity detection boxes; marking severe hallucination when no valid entity detection box is obtained, marking low-confidence local hallucination when the confidence of the entity detection box is lower than a preset threshold, or marking attribute mismatch hallucination when the entity attributes do not match the category or location of the entity detection box.

[0018] By adopting the above technical solution, the entities mentioned in the VLM response are independently verified through the visual positioning model. The three-level illusion classification provides a refined basis for subsequent quality scoring and optimization strategies.

[0019] Preferably, the response quality score includes: accuracy score, relevance score, and security score; The accuracy score is determined based on the hallucination detection results; the relevance score is determined by calculating the cross-modal similarity between the user request, the original image visual features, and the VLM response using a cross-modal model; and the security score is determined based on whether the VLM response contains features of prompt injection attacks, unde-identified sensitive entities, or preset violation categories. The accuracy score, relevance score, and security score are weighted and summed according to preset weights to obtain the response quality score. Based on image-aware hashing, a historical response quality score baseline for similar images is maintained. Responses that deviate from the historical response quality score baseline, or have hallucination detection anomalies, abnormal security detection results, or negative user feedback, are identified as samples to be optimized.

[0020] By adopting the above technical solutions, the multi-dimensional scoring system comprehensively evaluates the accuracy, relevance, and safety of VLM responses. Combined with deviation detection of historical baselines, abnormal responses can be identified and tracked in a timely manner.

[0021] Preferably, the user feedback data includes: Implicit feedback obtained through front-end interaction tracking points, business feedback obtained through downstream business metrics, and explicit ratings or negative reviews obtained through the user interface. The process of determining the image processing strategy to be optimized includes: associating and storing the hallucination type, preprocessing parameters, desensitized regions, response quality scores, and user feedback data in the sample to be optimized. For frequently occurring target categories, adjust the maximum long side parameter or the quality parameter of the format conversion for proportional non-scaled; Adjust the maximum long side, target format, quality parameters, or compression level parameters for preprocessing parameters that cause entity loss or blurred details; For samples with abnormal security scores, update the security restriction verification strategy and adjust the masking area or masking method based on samples with low relevance caused by masking.

[0022] By adopting the above technical solutions, user feedback and automated detection results are combined to form a quality closed loop of human-machine collaboration. Differentiated optimization strategies for different types of anomalies enable the system to continuously improve itself.

[0023] Preferably, the method further includes: performing failure degradation retry during the preprocessing stage; when the HTTP status code returned by the upstream VLM service is greater than or equal to 400 and failure retry is configured to be allowed, the original de-identified request body is resent to the upstream VLM service, and a degradation retry flag is marked in the metadata.

[0024] By adopting the above technical solution, the failure degradation retry mechanism automatically falls back to the original image for retry when VLM service is rejected due to preprocessing, thereby improving the robustness of the overall process.

[0025] Secondly, an image intelligent preprocessing and privacy desensitization system for VLM inference includes: The request parsing and image recognition unit is used to receive VLM inference requests, recursively traverse the request body and match data URLs, and identify Base64 encoded image data. A security verification unit is used to perform format whitelist verification and security restriction verification on the image data; The image preprocessing unit is used to perform proportional scaling without magnification and format conversion or re-encoding on the verified image data; A privacy desensitization unit is used to perform region masking on the preprocessed image and update the VLM inference request in response to the VLM inference request being an assertion-type request. The request sending unit is used to send the updated VLM inference request to the upstream VLM service and obtain a VLM response. The metadata recording unit is used to record image processing metadata and VLM response associated metadata. In scenarios where verification fails or the request is terminated, the skip status and skip reason are recorded in advance. The hallucination detection unit performs entity cross-validation on the VLM response to obtain hallucination detection results; the quality scoring unit generates a response quality score based on the hallucination detection results, semantic matching results, and security detection results; and the feedback optimization unit collects user feedback data and asynchronously performs hallucination detection, score calculation, feedback collection, and policy updates. Specifically, the processing strategies of the security verification unit, the image preprocessing unit, and the privacy desensitization unit are adjusted based on the image processing metadata, the hallucination detection results, the response quality score, and the user feedback data.

[0026] In summary, this application includes at least one of the following beneficial technical effects: 1. Improve image processing compatibility and transmission efficiency: By recursively traversing the request body to accurately extract image data and performing proportional scaling without enlargement and format conversion / re-encoding, image specifications are unified, avoiding request rejection or compatibility failures caused by high resolution or illegal formats. At the same time, the image size is significantly reduced, and network bandwidth and token quota consumption are reduced.

[0027] 2. Ensure data privacy and security: Automatically perform region masking and desensitization processing on assertion-type requests. Without affecting the integrity of the test functions, sensitive data such as account information and session identifiers in screenshots are masked, effectively eliminating the risk of privacy leakage when sending data to third-party VLM services and meeting data security compliance requirements.

[0028] 3. Achieve closed-loop optimization of response quality: Cross-validate the named entity recognition model with visual localization results to detect hallucinations, and generate a multi-dimensional response quality score by combining semantic matching and security detection. Identify samples to be optimized by collecting user feedback, and dynamically adjust the validation, preprocessing, and desensitization strategies to form a continuous optimization data closed loop of "detection -> scoring -> feedback -> strategy update -> re-evaluation".

[0029] 4. Enhance system fault tolerance and traceability: The non-blocking three-level security limit verification and failure degradation retry mechanism is adopted to avoid the preprocessing failure directly blocking the entire process, ensuring the continuity and robustness of the test process; at the same time, the image processing status (processed, no processing required, skipped) and associated metadata are recorded throughout the process, ensuring that every processing conclusion can be traced back to the specific image and response entity. Attached Figure Description

[0030] Figure 1This is a schematic diagram of the entire process and data closed loop of the intelligent image preprocessing and privacy desensitization method for VLM inference of the present invention; Figure 2 This is a schematic diagram of the recursive traversal and image recognition process of the present invention; Figure 3 This is a schematic diagram of the security verification and image preprocessing process of the present invention; Figure 4 This is a schematic diagram of the privacy desensitization and masking process of the present invention; Figure 5 This is a diagram of the image processing state machine of the present invention; Figure 6 This is a schematic diagram of the hallucination detection and response quality scoring process of the present invention; Figure 7 This is a schematic diagram of the image intelligent preprocessing and privacy desensitization system architecture of the present invention. Detailed Implementation

[0031] The following is in conjunction with the appendix Figure 1-7 This application will be described in further detail.

[0032] Example 1 This embodiment details the complete execution flow of an image intelligent preprocessing and privacy desensitization method for VLM inference. For example... Figure 1 As shown, the method in this embodiment includes VLM inference request reception, image recursive recognition, multi-level security restriction verification, image preprocessing, assertion-type request recognition and privacy desensitization, request sending, failure degradation retry, metadata recording, hallucination detection, response quality scoring, user feedback collection and strategy optimization, and forms a data closed loop through subsequent re-evaluation of similar requests; including the following steps: Step S1: Receive VLM inference request The test execution engine generates VLM inference requests during the automated testing process. These requests are used for operations such as UI element identification and assertion judgment. The request body is JSON structured data, which contains one or more Base64 encoded screenshots of the UI.

[0033] Step S2: Recursive Traversal and Image Recognition like Figure 2 As shown, in step S2, a recursive traversal is performed according to the data structure type of the request body node; for string nodes, image Data URL matching is further performed to extract MIME type, Base64 data, request path and image index, and a recursive traversal is performed on the request body of the VLM inference request to extract all image data.

[0034] The specific execution process is as follows: The request body is traversed using a depth-first search based on its data structure type. When the value encountered is an object, a new object is created, and the traversal function is recursively called for each attribute. When the value encountered is an array, the traversal function is recursively called for each element. When the value is a string, it is attempted to be parsed as a Data URL. Values ​​that are not objects, arrays, or strings are returned as is. During the traversal, an incrementing image index counter is maintained to ensure that image numbers across nested levels are unique and ordered.

[0035] Data URL parsing uses the regular expression: ` / ^data:(image\ / (?:png|jpe?g|webp));base64,([a-zA-Z0-9+ / =\s-]+)$ / iu` for matching. The `i` flag makes the match case-insensitive, and the `u` flag enables full Unicode support. The regular expression captures both the image's MIME type and the Base64 encoded data. Only strings that match this regular expression are recognized as valid image data.

[0036] Output: A list of extracted image data, each item containing the image index, MIME type, Base64 encoded string, and path to the request body.

[0037] Step S3: Format whitelist verification The MIME type parsed in step S2 is matched against a preset list of valid formats. The preset list of valid formats includes: image / png, image / jpeg, image / jpg, and image / webp. Only image data that successfully matches is allowed to proceed to subsequent security restriction checks. Image data whose MIME type is not in the whitelist is directly marked as skipped, and the reason for skipping is recorded as type_unsupported.

[0038] Output: A list of image data that passed the format whitelist validation and the image records that were skipped.

[0039] Step S4: Security Restriction Verification Reference Figure 3 For image data that passes the format whitelist verification, three levels of security restriction verification are performed sequentially, with each level being non-blocking: Level 1: Image count limit per request: Counts the total number of images that pass the format validation in the current request. When the number exceeds the maximum number of images allowed in a single request (default value is 20), the excess images are marked as skipped, and the reason for skipping is recorded as image_count_limit.

[0040] Level 2: Byte limit after decoding a single image: Decode the Base64 data of each image. When the length of the decoded data exceeds 16MB, mark the image as skipped and record the reason for skipping as decoded_byte_limit.

[0041] Level 3: Single Image Pixel Limit: The image processing library reads the image dimensions and calculates the total number of pixels (width × height). When the total number of pixels exceeds 20 million, the image is marked as skipped, and the reason for skipping is recorded as pixel_limit. The image processing library's limitInputPixels parameter is set to 20 million, and the library automatically checks for pixel limits when reading images.

[0042] Image data that passes all three levels of verification is entered into step S5 as images to be preprocessed. Images that fail verification are marked as skipped, with skipping reasons supported by image_count_limit (exceeding the limit for the number of images), decoded_byte_limit (exceeding the limit for the number of decoded bytes), pixel_limit (exceeding the limit for the total number of pixels), or processing_failed (processing failed).

[0043] Output: A list of images to be preprocessed that have passed security limit checks, along with records of skipped steps at each level.

[0044] Step S5: Scaling without enlarging The image to be preprocessed is scaled proportionally without magnification, using the longer side of the image as a reference.

[0045] The specific execution process is as follows: The width and height of the original image are read using the `metadata()` method of the image processing library. If reading the dimensions fails, the reason for skipping is recorded as `missing_dimensions`, and the original values ​​are retained without further processing.

[0046] The judgment is based on the maximum long side parameter (default value is 1280 pixels): If the longer side of the original image (the larger of the width and height) does not exceed the maximum longer side, no scaling is performed, and the target size is equal to the original size (no enlargement principle). When the longer side of the original image exceeds the maximum longer side, the aspect ratio is maintained during scaling. The calculation method is as follows: if the width is the longer side, the target width equals the maximum longer side, and the target height equals the original height multiplied by (maximum longer side divided by original width) and rounded to the nearest integer; if the height is the longer side, the target height equals the maximum longer side, and the target width equals the original width multiplied by (maximum longer side divided by original height) and rounded to the nearest integer. Ensure that the minimum value of the target width and target height is 1.

[0047] To clearly illustrate the scaling calculation logic of this scheme, which uses the longer side as the reference, maintains the aspect ratio, and does not enlarge the image, size conversion examples are given using multiple sets of original images at different resolutions. Specific data is shown in Table 1: Table 1. Example of scaling calculation without enlargement. Step S6: Preprocessing mode decision and execution coding like Figure 3 As shown, the identified image data first undergoes format whitelist verification and security restriction verification. Images that pass the verification are used as images to be preprocessed, and are then scaled proportionally without enlargement and re-encoded according to the original size, maximum long side, target format, and quality parameters. The preprocessing mode decision is based on the mode parameters configured in the settings. Adaptive mode (auto): Scaling and re-encoding are performed only when the image needs to be scaled; if the image does not need to be scaled and the source format is the same as the target format, re-encoding is not performed, and the original data remains unchanged.

[0048] Fixed mode: Forces re-encoding regardless of whether the image needs to be scaled.

[0049] The mode decision also considers the differences between the source and target formats: when scaling is required but the source and target formats are different, re-encoding will be performed even in adaptive mode.

[0050] Scaling uses the resize() method from the image processing library, employing the fit:inside mode to maintain the proportions of the inner frame, and setting the non-enlargement flag (withoutEnlargement:true).

[0051] Encoding uses target format and quality parameters: (1) WebP format: use webp({quality}), the default value of the quality parameter is 80; (2) JPEG format: use jpeg({quality}), the default value of the quality parameter is 85; (3) PNG format: Use png({compressionLevel:9}), the default value of the compression level parameter is 9.

[0052] The encoded output is a buffer, and the number of bytes is recalculated. If both the original and target formats are PNG and no scaling is required, the original buffer remains unchanged. Other output formats are uniformly converted to JPEG or WebP. When the original format is PNG containing an alpha channel and the target format is JPEG, the alpha channel is first composited with a preset background color (such as white or black) before conversion to avoid loss of transparent area information affecting VLM recognition.

[0053] Output: Preprocessed image buffer, processed format, processed number of bytes, and quality parameters.

[0054] This solution supports both adaptive and fixed preprocessing modes, and can complete image format transcoding and compression as needed. Different original images have significant differences in size, number of bytes, and compression effect after encoding. Typical conversion data are shown in Table 2. Table 2 Examples of Format Conversion and Re-encoding Results Step S7: Assertion Request Identification and Region Masking Processing like Figure 4 As shown, when a VLM inference request is identified as an assertion request, the masking area is determined according to the specified region masking strategy or the bottom ratio masking strategy, and the masked and desensitized image is written back to the VLM inference request to determine whether the current VLM inference request is an assertion request.

[0055] Assertion-type requests are identified in at least one of the following ways: Identification is based on the task step type carried in the request: when the task step type field value is an assertion type (such as assertion, verify, etc.), it is determined to be an assertion type request; Identify assertion intent based on the request context: Analyze the metadata fields in the request context, and determine that the request is an assertion-related request if it contains assertion-related identifiers; The call intent is identified based on the analysis of the request body; semantic analysis is performed on the text content in the request body, and when intents such as assertion, verification, and validation are identified, it is determined to be an assertion request.

[0056] In response to the identification that the VLM inference request is an assertion-type request, region masking processing is performed. Masking region calculation supports two strategies: Specified area masking strategy: The user provides coordinates and dimensions. A transformation function processes the normalized values ​​(multiplying the image dimensions within the range of 0-1 to obtain pixel coordinates) and pixel values ​​(values ​​greater than 1 are directly used as pixel coordinates). Boundary cropping is applied to the coordinates and dimensions to ensure that the masked area does not exceed the image boundaries.

[0057] Bottom Proportion Masking Strategy: Users provide bottom masking proportion parameters (e.g., 0.15 represents the bottom 15% area). The width of the masking area is equal to the width of the image, and the height of the masking area is equal to the height of the image multiplied by the proportion value and rounded to the nearest integer. The position of the masking area extends upward from the bottom of the image.

[0058] Step S8: Mask Compositing Execution Create an overlay image with a width and height equal to the masking area, 4 channels (RGBA), and a pure black opaque fill (r=0, g=0, b=0, alpha=1). Use the composite() method of an image processing library to composite the overlay image onto the preprocessed image with coordinates top=y-coordinate of the masking area and left=x-coordinate of the masking area. Convert the composite image to a buffer, then to a Base64 string to obtain the de-identified image.

[0059] Replace the original image data with an anonymized image and update the image field in the VLM inference request.

[0060] Output: De-identified VLM inference request and metadata of the masked region (masking strategy type, masking coordinates, masking size).

[0061] Step S9: Sending a request and receiving a VLM response The updated VLM inference request is sent to the upstream VLM service to obtain a VLM response.

[0062] Failure-based retry mechanism: In the VLM proxy request processing flow, after the preprocessing stage is completed and the request is sent, if the upstream VLM service returns an HTTP status code greater than or equal to 400 and failure retry is configured to allow, the error message is first parsed. If the error message indicates that the image size or total number of pixels exceeds the limit, a failure-based retry is not performed, and the request is directly recorded as failed. If the error message is not due to size exceeding the limit, a failure-based retry is triggered: the original request body (which has undergone privacy masking but not scaling and format conversion) is resent to the upstream VLM service, the failure-based retry flag is set to true in the metadata, and the original preprocessed metadata and the failure-based retry metadata are marked. A retry response is returned instead of the initial failure response.

[0063] Output: VLM response content and HTTP status code.

[0064] Step S10: Metadata Recording Record the associated metadata corresponding to each image, including: Image processing metadata: original format, processed format, original width, original height, processed width, processed height, original number of bytes, processed number of bytes, quality parameters; such as Figure 5As shown, each image enters only one mutually exclusive state—processed, no processing required, or skipped—during a single processing cycle, and the state information and related metadata are written into the record results after the state is determined.

[0065] Processing status: One of three states: processed, unnecessary, or skipped; Skip reason: Record the specific reason when skipping (type_unsupported, image_count_limit, decoded_byte_limit, pixel_limit, missing_dimensions, processing_failed, etc.); VLM response associated metadata: response timestamp, response length, and downgrade retry flag.

[0066] The state machine ensures that each image has one and only one state, and that the state information is fully recorded in the metadata.

[0067] Output: Complete associated metadata records. The system collects complete associated metadata for each input image throughout the entire process, distinguishing between three processing states: processed, no processing required, and skipped, and records the corresponding parameters and reasons for any anomalies. A sample metadata storage format is shown in Table 3. Table 3. Example of Image Processing Metadata Record Step S11: Hallucination detection and correlation of detection results like Figure 6 As shown, response entities and attributes are extracted from the VLM response, and cross-validation is performed using the visual localization results of the original image to determine severe hallucinations, low-confidence local hallucinations, or attribute mismatch hallucinations. Subsequently, a response quality score is generated based on accuracy, relevance, and security. After acquiring the VLM response, hallucination detection does not simply perform a literal judgment on the response text; instead, it remaps the entity attributes declared in the response text to the original image retained before sending, and performs cross-validation using the visual localization results. Here, the original image refers to the image before scaling, re-encoding, and region masking; this avoids interference from masking regions, compression artifacts, or size changes in determining whether entities actually exist.

[0068] Specifically, the metadata recording unit generates a request association identifier for each VLM inference and associates and saves the request association identifier with the image index, original image, preprocessed image, desensitized image, processing status, masking parameters, and VLM response. The hallucination detection unit reads the request association identifier to obtain the original image and image processing context that correspond one-to-one with the current VLM response, thereby ensuring that each subsequent detection conclusion can be traced back to a specific image and a specific response entity.

[0069] In this embodiment, hallucination detection includes the following sub-steps: Step S111: Extract the response entity and its attributes.

[0070] A Named Entity Recognition (NER) model is used to parse the VLM response and generate entity records according to the order of description in the response. The NER model can perform sequence labeling and semantic understanding on the unstructured natural language text generated by VLM, accurately identify and extract entity words with specific meanings, and map them into structured attribute data.

[0071] In a preferred embodiment, the named entity recognition model employs an architecture based on a pre-trained language model (such as BERT or RoBERTa) combined with a Conditional Random Field (CRF). The model receives a character sequence of the VLM response text as input and outputs a labeled sequence with entity boundaries and categories. The labeling system adopts the BIOES or BIO2 labeling specifications and defines dedicated entity type labels adapted to the image description scenario, including but not limited to: interface elements (such as buttons, input boxes, icons, etc.), quantity, color, and spatial location (such as top left corner, bottom, etc.).

[0072] During the extraction process, the model first identifies the core object category entity (such as "button") in the text. Then, through dependency parsing or slot filling based on an attention mechanism, it binds modifiers scattered around the core entity (such as "red," "1," and "bottom right corner") to that core entity, forming a complete entity record. Each entity record includes at least an entity identifier and an object category, and may further include at least one entity attribute from quantity, color, and spatial location. For example, for the response text "There is a red button in the bottom right corner of the image," the named entity recognition model not only identifies the object category entity "button," but also associates "one" as a quantity attribute, "red" as a color attribute, and "bottom right corner" as a spatial location attribute, thus combining them to generate a complete entity record. If the VLM response contains multiple entities, multiple entity records are created separately to avoid mixing and comparing different entity categories, quantities, colors, or spatial locations.

[0073] Output: A list of response entities corresponding to the request-associated identifier; each response entity has structured entity attributes that can be verified by the visual positioning model.

[0074] Step S112: Detection task construction and visual localization.

[0075] For each responding entity, its object category is used as the target to be localized. When the responding entity also carries quantity, color, or spatial location attributes, these attributes are also written into the entity's detection task. The detection task and the original image are input into the visual localization model to obtain one or more entity detection boxes associated with the entity. Each entity detection box carries at least the detection category, detection box location, and detection confidence. When multiple candidate detection boxes exist, the detection boxes that match the entity attributes are retained as valid candidate detection boxes.

[0076] For the quantity attribute, the number of valid candidate detection boxes that satisfy the constraints of object category and other available attributes is counted. For the spatial location attribute, the region where the detection box is located is determined based on its position in the original image and compared with the spatial location of the response entity. For the color attribute, the corresponding color features are determined based on the image content of the area covered by the detection box and compared with the color attribute of the response entity. Thus, an entity description in the response text is transformed into a detection task that can be verified on the original image.

[0077] Output: The set of entity detection boxes corresponding to each response entity, detection confidence, quantity comparison results, color comparison results, and spatial location comparison results.

[0078] Step S113: Hallucination type determination.

[0079] Based on the output of step S112, generate hallucination detection results according to preset judgment rules: When no valid entity detection box corresponding to the response entity is obtained, it means that the entity in the response cannot be visually located in the original image, and the entity is marked as a severe hallucination. When an entity detection bounding box is obtained but its detection confidence is lower than a preset threshold, the entity is marked as a low-confidence local illusion; the preset threshold can be configured by the system and can be set according to different image categories or calling scenarios; When an entity detection box has located the corresponding entity, but the entity's category, quantity, color, or spatial location does not match the detection result, the entity is marked as an attribute mismatch illusion. When all verifiable attributes of the responding entity are consistent with the detection results, the entity is marked as an undetected hallucination.

[0080] When a single VLM response contains multiple response entities, hallucination detection results are generated for each response entity. Subsequently, a summary hallucination detection result for the VLM response is formed based on the detection results of each entity. The summary hallucination detection result records at least the total number of response entities, the entity identifier corresponding to each hallucination type, the entity detection box, and the reason for the mismatch, so that the quality scoring unit and the feedback optimization unit can use the same entity-level evidence.

[0081] Output: Hallucination detection results, including entity-level detection results and response-level summary results. Entity-level detection results include at least the entity identifier, entity attributes, entity bounding box, detection confidence, hallucination type, and reason for mismatch.

[0082] The hallucination detection module independently performs visual cross-validation on each entity extracted from the VLM response. Based on the detection bounding box, confidence score, and attribute matching results returned by visual localization, it classifies the hallucination type. Entity-level judgment examples are shown in Table 4: Table 4. Example of the hallucination detection process Step S12: Response quality scoring and scoring basis recording After obtaining the hallucination detection results, the quality scoring unit generates a response quality score based on three dimensions: accuracy, relevance, and security. The inputs for all three dimensions are bound to the same request-associated identifier; therefore, any comprehensive score can be traced back to the corresponding user request, original image, VLM response, entity-level hallucination detection results, and security detection results.

[0083] Step S121: Accuracy score generation.

[0084] The quality scoring unit reads the summarized hallucination detection results at the response level and determines the accuracy score according to a preset hallucination type-score mapping rule. For response entities where no hallucination is detected, a higher accuracy rating is assigned. For severe hallucinations, low-confidence partial hallucinations, and attribute mismatch hallucinations, accuracy ratings are generated according to their respective deduction rules. If multiple response entities exist for the same response, the accuracy ratings of each entity are combined into a response-level accuracy score according to a preset summarization rule. This summarization rule allows for a higher deduction level for severe hallucinations compared to low-confidence partial hallucinations or attribute mismatch hallucinations, thus ensuring that the accuracy score reflects the differences in the impact of different hallucination types.

[0085] Step S122: Generate correlation score.

[0086] The quality scoring unit uses user requests, original image visual features, and VLM responses as the objects of relevance calculation. Specifically, it obtains the semantic features of the user request, the visual features of the original image, and the semantic features of the VLM response, respectively. It then calculates the cross-modal semantic matching degree between the VLM response and the user request, and between the VLM response and the visual features of the original image. Finally, it obtains a relevance score according to a preset fusion rule. Using the original image instead of anonymized images in this scoring allows the system to simultaneously identify two types of situations: "response does not match user request" and "anonymization or preprocessing leads to the omission of key image information in the response."

[0087] Step S123: Security score generation.

[0088] The quality scoring unit performs security checks on the VLM response to determine whether it contains features indicating injection attacks, unmasked sensitive entity features, or preset violation categories. The security detection result indicates at least whether the corresponding security feature was detected and the type of the detected security feature. The quality scoring unit generates a security score based on the security detection result and preset security scoring rules; when a security feature is detected, the security score is reduced according to the corresponding rules, and the type of security feature that triggered the reduction is written into the associated metadata.

[0089] Step S124: Generate comprehensive score.

[0090] The accuracy score, relevance score, and security score are multiplied by their respective preset weights and then summed to obtain the response quality score. The accuracy score is used as the weighting factor. The correlation score is Safety score: And its corresponding weight is , and For example, response quality score satisfy: Response quality score ; in, , and The weights are pre-configured by the system, and the sum of the three is 1. These weights can be configured for different calling scenarios; for example, the weight corresponding to the accuracy score can be increased for assertion-type requests, and the weight corresponding to the security score can be increased for calling scenarios with higher privacy risks. The weight adjustments take effect in subsequent scoring and are saved in association with the configuration version for the corresponding scenario.

[0091] Output: Response quality score and its scoring criteria, which include at least accuracy score, relevance score, safety score, weights of each dimension, hallucination detection result identifier, and safety detection result identifier.

[0092] The response quality score is calculated by weighting three dimensions: accuracy, relevance, and security. Each score result can be fully traced back to the underlying basis such as hallucination detection and security detection. Examples of score calculations for multiple requests are shown in Table 5: Table 5. Examples of Response Quality Scores and Scoring Criteria Step S13: Baseline maintenance of similar images and identification of samples to be optimized To transform a single scoring result into a data loop that can be used for optimization, the system establishes closed-loop sample records based on the requested association identifier. Each closed-loop sample record is associated with at least: image perceptual hash, image processing metadata, masking strategy and masking parameters, VLM response, entity-level and response-level illusion detection results, accuracy score, relevance score, security score, response quality score, and user feedback data.

[0093] For each image, a perceptual hash is calculated. Specifically, a perceptual hashing algorithm (such as pHash, aHash, or dHash) is used to reduce the image to a fixed size (e.g., 8×8 pixels, 64 pixels in total), convert it to a grayscale image, and calculate the average grayscale value of all pixels. The grayscale value of each pixel is compared with the average value; a value greater than or equal to the average is recorded as 1, and a value less than is recorded as 0, thus generating a 64-bit binary hash sequence, which serves as the fingerprint of the image. When classifying images into the same category, the Hamming distance (i.e., the number of different characters at corresponding positions) between the current image hash sequence and historical image hash sequences is calculated. When the Hamming distance is less than a preset similarity threshold (e.g., less than or equal to 5), the images are classified as images of the same category. Simultaneously, images with the same image processing parameters (e.g., maximum long side, target format, quality parameters, etc.) are also classified as images of the same category. For images of the same category, a historical response quality score baseline is maintained. The baseline is dynamically and smoothly updated based on the recorded historical response quality scores using a sliding window or exponentially weighted moving average (EWMA) method, and is associated with image category, calling scenario, and configuration version, thereby avoiding direct mixing and comparison of scores from different types of images, different calling scenarios, or different preprocessing configurations.

[0094] For images of the same type, a historical response quality score baseline is maintained. This baseline can be determined based on the recorded historical response quality scores and associated with the image category, calling scenario, and configuration version, thereby avoiding direct comparison of scores from different image types, calling scenarios, or preprocessing configurations.

[0095] Subsequently, the system reads the response quality score of the current closed-loop sample and compares it with the historical response quality score baseline of the corresponding similar images. If the current score deviates from the baseline, or if the current closed-loop sample has hallucination detection anomalies, abnormal security detection results, or negative user feedback, then the closed-loop sample is marked as a sample to be optimized according to preset sample identification rules. The sample to be optimized does not only retain the score value, but also retains the entity-level detection evidence, image processing parameters, masking parameters, and user feedback that caused the score, so that subsequent optimization can be located at specific processing stages.

[0096] Output: A list of samples to be optimized. Each sample to be optimized includes at least the request association identifier, image perceptual hash, anomaly type, score deviation information, associated metadata, hallucination detection result, response quality score, and user feedback data.

[0097] Step S14: Optimization of Differentiation Strategy Based on Closed-Loop Samples After receiving the sample to be optimized, the feedback optimization unit does not directly modify the original image or historical records. Instead, it generates a corresponding candidate optimization strategy based on the anomaly type and invokes this candidate optimization strategy in subsequent VLM inference requests for the same or similar images. The candidate optimization strategy is associated with its source sample to be optimized, the applicable image-aware hash category, the invocation scenario, and the configuration version, so as to track the optimization results after a response quality score is generated again.

[0098] Specifically, the following processes are included: Step S141: Policy update for security anomaly samples.

[0099] When the security score of a sample to be optimized is abnormal, the feedback optimization unit reads the security detection result, image processing metadata, and current security restriction verification policy associated with that sample. Based on the security feature type in the security detection result, the security restriction verification policy is updated, ensuring that subsequent requests use the updated policy during format whitelist verification, security restriction verification, or related security restriction rule execution phases. The updated security restriction verification policy is saved as a new policy version and associated with the sample to be optimized that triggered the update. Thus, when the same or similar images re-enter the system, they can be processed according to the updated security restriction verification policy before being sent to the VLM service.

[0100] Step S142: Masking optimization for low-relevance samples caused by hallucination.

[0101] When the sample to be optimized simultaneously presents hallucination detection results and low correlation scores, the feedback optimization unit reads the entity-level detection results, the position of the entity detection box in the original image, the current masking region, and the masking method. If the image content indicated by the entity detection box is related to the current masking region, the coordinates, size, or bottom ratio of the masking region is adjusted according to the entity detection box, the current masking region, and privacy desensitization requirements. If the masking region processing method is not suitable for the current image category, the masking method is adjusted between specified region masking and bottom ratio masking. The adjusted masking parameters still need to be cropped at the boundaries and saved as a new masking configuration version.

[0102] This process does not negate privacy masking; rather, it ensures that subsequent VLM inference requests retain image information relevant to the user request and not belonging to the protected area, while maintaining the masking constraints. The adjusted masking configuration is associated with the image-aware hash category so that it can be invoked in subsequent requests for the same or similar images.

[0103] Step S143: Optimization result write-back and re-evaluation.

[0104] When a subsequent VLM inference request enters the preprocessing stage, the system first calculates the image-aware hash of the currently requested image and retrieves historical candidate optimization strategies from the policy version library that meet a preset threshold in terms of Hamming distance. If a match is found, the system automatically loads the corresponding policy version (such as adjusted maximum long side, quality parameters, masking region coordinates, etc.) to perform compensation processing on the request, and writes the policy version used this time into a new closed-loop sample record. After the VLM response is returned, steps S11 to S13 are repeated to obtain the hallucination detection result, response quality score, and user feedback data again. The new score result is compared with the historical response quality score baseline before optimization and the source sample to be optimized: if the new score result meets the preset improvement judgment rule, the candidate optimization strategy is retained as the current effective strategy and marked as a stable version; if it does not meet the improvement judgment rule or the score has decreased compared to before, a rollback mechanism is triggered, the currently loaded compensation strategy is canceled and the system's default preprocessing and desensitization parameters are restored, and the result is retained as a new sample to be optimized, and the policy optimization process is re-entered according to the new anomaly type. This forms a closed data loop: "Image and processing metadata → VLM response → Illusion detection → Quality scoring → Identification of samples to be optimized → Policy update → Accurate compensation through hash matching → Re-evaluation of subsequent requests." Each stage in the loop is associated with a request-related identifier, image-aware hash, and configuration version, which allows for tracing the source of evidence for a policy adjustment and verifying the effectiveness of that policy adjustment in subsequent similar image requests.

[0105] Output: Updated security restriction verification policy or masking configuration, corresponding configuration version, policy source, sample identifier to be optimized, and subsequent re-evaluation results.

[0106] The implementation principle of this embodiment of an intelligent image preprocessing and privacy desensitization method for VLM inference is as follows: By deploying a preprocessing and privacy desensitization middleware layer before sending the VLM inference request, the image data in the request body is recursively extracted, security verified, size optimized, and formatted. Sensitive area masking is automatically performed in assertion scenarios, thereby eliminating the risk of privacy leakage while ensuring the integrity of the testing function. After the preprocessed image is sent to the VLM service, the output quality is controlled through illusion detection and response quality scoring, forming a closed loop of continuous optimization combined with user feedback. The non-blocking skip mechanism and failure degradation retry ensure the robustness of the process, and the full metadata recording realizes the traceability of the processing process.

[0107] Example 2 Reference Figure 1 The difference between this embodiment and embodiment 1 is that it uses differentiated configuration parameters and processing strategies for different calling scenarios.

[0108] The system architecture of this invention comprises one core service and three invocation scenarios. The core service is a preprocessing service, providing three public methods: `preprocess()` for complete preprocessing of VLM proxy requests, `preprocessGeneration()` for performing strict validation preprocessing for component image generation, and a dedicated preprocessing method for document image extraction. The three scenarios share the same service instance through dependency injection.

[0109] (1) VLM agent Read preprocessing parameters from the project test configuration, which includes enable flags, mode (adaptive or fixed), maximum long side, format, quality, and failure retry switch. Perform a complete recursive traversal, security checks, proportional scaling without zooming, format conversion, privacy masking, and failure degradation retry process.

[0110] This invention supports differentiated configurations for multiple business scenarios. Among them, the automated testing VLM proxy scenario focuses on process fault tolerance and non-blocking processing. The complete set of preprocessing and security verification parameter configuration examples for this scenario are shown in Table 6: Table 6 Examples of VLM agent scenario configuration parameters (2) Execution component image generation Using the default configuration (maximum long side 1280, WebP format, quality 80), a strict validation mode is executed. In this mode: exceeding the image quantity limit directly throws a business exception (error code VLM_IMAGE_COUNT_LIMIT); not supporting the image type directly throws an exception (error code VLM_IMAGE_TYPE_UNSUPPORTED); exceeding the byte limit directly throws an exception (error code VLM_IMAGE_SIZE_LIMIT); exceeding the pixel limit directly throws an exception (error code VLM_IMAGE_PIXEL_LIMIT); and failing to read the size directly throws an exception (error code VLM_IMAGE_DIMENSION_ERROR). Unlike the non-blocking skipping in the VLM proxy scenario, any validation failure in the generation scenario directly throws the corresponding exception, blocking subsequent processing.

[0111] (3) Extraction of images from requirements documents This uses a dedicated configuration for document preprocessing and supports concurrency control parameters. The primary scenario involves extracting images from a requirement document and sending them to VLM for document comprehension. Configuration parameters can be adjusted based on the typical characteristics of the document images.

[0112] The implementation principle of this embodiment is as follows: by providing differentiated configuration parameters and error handling strategies for different calling scenarios, scenario adaptation is achieved under the premise of sharing the core preprocessing logic. The VLM proxy scenario focuses on fault tolerance and non-blocking processing, the execution component image generation scenario focuses on strict verification and fast failure, and the document image extraction scenario focuses on concurrent processing capabilities and document feature adaptation.

[0113] Example 3 Reference Figure 7 This embodiment discloses an intelligent image preprocessing and privacy desensitization system for VLM inference, comprising the following functional units: The request parsing and image recognition unit is configured to receive VLM inference requests, recursively traverse the request body and perform data URL matching to identify Base64 encoded image data. This unit performs a depth-first traversal, maintains an incrementing image index counter, and uses the regular expression ` / ^data:(image\ / (?:png|jpe?g|webp));base64,([a-zA-Z0-9+ / =\s-]+)$ / iu` to perform Data URL matching on string values, capturing MIME types and Base64 encoded data.

[0114] The security verification unit is configured to perform format whitelist verification and security restriction verification on the image data. Format whitelist verification matches the MIME identifier of the image data against a preset list of valid formats (image / png, image / jpeg, image / jpg, image / webp). Security restriction verification sequentially applies three levels of restrictions: a limit on the number of images requested per request (default 20), a limit on the number of bytes after decoding a single image (default 16MB), and a limit on the total number of pixels in a single image (default 20 million). Images that fail verification are marked as skipped, and the reason for skipping is recorded.

[0115] The image preprocessing unit is configured to perform proportional scaling without enlargement and format conversion or re-encoding on verified image data. It maintains the aspect ratio when the original image's long side exceeds the maximum long side parameter (default 1280 pixels). During encoding, it generates preprocessed results according to the target format (default WebP) and quality parameters (default 80) or compression level parameters (PNG default 9). It supports two preprocessing modes: adaptive (auto) and fixed (fixed).

[0116] The privacy desensitization unit is configured to perform region masking processing on the preprocessed image and update the VLM inference request in response to the VLM inference request being an assertion-type request. Assertion-type requests are identified through task step type, contextual assertion intent, or invocation intent obtained based on request body analysis. Region masking processing supports two strategies: specified region masking (user-provided coordinates and dimensions) and bottom-scale masking (user-provided bottom-scale parameters). Masking coordinates support both normalized scale values ​​within the closed interval [0,1] and direct use of pixel coordinate values, and the masked region is cropped at its boundaries.

[0117] The request sending unit is configured to send the updated VLM inference request to the upstream VLM service and obtain a VLM response. When the upstream VLM service returns an HTTP status code greater than or equal to 400 and is configured to allow retries on failure, a failure degradation retry is triggered, and the original, de-identified request body is resent to the upstream VLM service.

[0118] The metadata recording unit is configured to record image processing metadata and VLM response-related metadata. In scenarios where validation fails or the request terminates, the skip status and skip reason are recorded in advance. Each image goes through one of three states: processed, unnecessary, or skipped. The state machine ensures that each image has exactly one state.

[0119] The hallucination detection unit is configured to extract response entities based on the VLM response and cross-validate them with the visual localization results in the original image. It uses a named entity recognition model to extract at least one entity attribute from entity category, quantity, color, and spatial location. These entity attributes are then used to construct a detection task, which is input into the visual localization model to obtain entity detection boxes. Based on the detection results, the unit labels the hallucination as severe hallucination, low-confidence local hallucination, or attribute mismatch hallucination.

[0120] The quality scoring unit is configured to generate a response quality score based on hallucination detection results, cross-modal similarity, and security detection results. The response quality score includes an accuracy score, a relevance score, and a security score, which are weighted and summed according to preset weights (accuracy 0.4, relevance 0.35, security 0.25). A historical response quality score baseline for similar images is maintained based on image-aware hashing, and samples deviating from the baseline are identified for optimization.

[0121] The feedback optimization unit is configured to collect user feedback data (implicit feedback obtained through front-end interaction tracking points, business feedback obtained through downstream business metrics, and explicit ratings or negative reviews obtained through the user interface), identify samples to be optimized based on associated metadata, hallucination detection results, response quality scores, and user feedback data, update security restriction verification strategies for samples with abnormal security scores, and adjust the masking area or masking method for low-relevance samples caused by hallucinations.

[0122] The aforementioned functional units work collaboratively to form a complete processing chain that establishes a closed-loop quality control system from request reception to response. The request parsing and image recognition unit extracts image data and then passes it to the security verification unit for filtering. Data that passes verification is optimized for size and format by the image preprocessing unit. For assertion-type requests, the privacy desensitization unit performs region masking before the request sending unit sends it to the VLM service. The metadata recording unit records data at each stage throughout the entire process. After the VLM response is returned, the illusion detection unit and quality scoring unit perform quality assessment, and the feedback optimization unit continuously optimizes the system strategy based on the assessment results and user feedback.

[0123] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. An intelligent image preprocessing and privacy desensitization method for VLM inference, characterized in that, include: Receive VLM inference requests; The request body of the VLM inference request is recursively traversed, and the strings obtained by the traversal are matched with data URLs to identify the Base64 encoded image data. The identified image data is subjected to format whitelist verification and security restriction verification, and the image data that passes the verification is used as the image to be preprocessed. If the verification fails and the request is terminated, record the corresponding skip status and skip reason metadata in advance; Using the long side of the image as a reference, the image to be preprocessed is subjected to proportional scaling without enlargement and format conversion or re-encoding to obtain a preprocessed image; In response to the identification that the VLM inference request is an assertion request, region masking processing is performed on the preprocessed image based on the masking region to obtain a desensitized image, and the VLM inference request is updated using the desensitized image. Send the updated VLM inference request to the upstream VLM service and obtain a VLM response; Record the original size, processed size, number of bytes, processing parameters, verification status, skipping reason, and associated metadata of the VLM response corresponding to each image; The response entity is extracted based on the VLM response, and the response entity is cross-validated with the visual localization result in the original image to obtain the hallucination detection result. Based on the hallucination detection results, the semantic matching results between the VLM response and the user request, and the security detection results of the VLM response, a response quality score is generated. Collect user feedback data associated with the VLM response; Based on the associated metadata, the hallucination detection results, the response quality score, and the user feedback data, samples to be optimized are determined according to the dimensions of the same image perceptual hash or the same image processing parameters. Based on the samples to be optimized, at least one of the following is adjusted: the verification strategy of the format whitelist verification and security restriction verification, the preprocessing parameters of the proportional scaling without enlargement and format conversion or re-encoding, and the desensitization strategy of the region masking processing. The adjusted processing strategy is associated with the corresponding image processing features so that subsequent VLM inference requests with the corresponding image processing features adopt the optimized intelligent image preprocessing and privacy desensitization strategy.

2. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The recursive traversal includes: performing a depth-first or breadth-first traversal on the multi-level nested request body based on the data structure type to extract image data; The format whitelist verification includes: identifying the current image format by parsing the file header features or multimedia type identifier of the image data, matching the current image format with a preset legal format feature library, and only allowing successfully matched image data to enter the security restriction verification.

3. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The security restriction verification includes: sequentially executing the image quantity limit, the byte limit after decoding a single image, and the total pixel limit of a single image; For image data that fails the security limit verification, it is marked as skipped and the skipping reason is recorded, such as exceeding the limit for the number of images, the limit for the number of decoded bytes, the limit for the total number of pixels, or the processing failure.

4. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The proportional scaling without enlargement includes: maintaining the aspect ratio during scaling when the longer side of the original image is greater than a preset maximum longer side; the format conversion or re-encoding includes: A preprocessed image is generated according to the quality parameters or compression level parameters corresponding to the target format; the preprocessing parameters include at least one of the following: maximum long side, target format, quality parameters, and compression level parameters.

5. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The VLM inference request was identified as an assertion request, including: Identify based on the task step type carried in the request, the assertion intent field in the request context, or the invocation intent obtained from the request body analysis; The region masking process includes specified region masking or bottom ratio masking; wherein, for masking coordinate values ​​located in the closed interval of 0 to 1, they are used as normalized ratio values ​​and converted into pixel coordinates; for masking coordinate values ​​greater than 1, they are used as pixel coordinates, and the resulting masking region is cropped at the boundary.

6. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, When the upstream VLM service returns an HTTP status code greater than or equal to 400 for the updated VLM inference request and is configured to allow retry on failure, the VLM inference request is reconstructed using the original image, the proportional scaling without zooming and the format conversion or re-encoding process are removed, and when the VLM inference request is an assertion-type request, the region masking process of the original image is retained. The reconstructed VLM inference request is sent to the upstream VLM service.

7. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The response entity is extracted from the VLM response, and the response entity is cross-validated with the visual localization result in the original image to obtain the hallucination detection result, including: extracting at least one entity attribute from the VLM response, including object category, quantity, color and spatial location, using a named entity recognition model; The entity attributes are constructed into a detection task and input into a visual localization model to obtain entity detection boxes in the original image; Mark severe illusion when no valid entity detection box is obtained, mark low-confidence local illusion when the confidence of the entity detection box is lower than a preset threshold, or mark attribute mismatch illusion when the entity attributes do not match the quantity, color or position corresponding to the entity detection box.

8. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The response quality score includes: accuracy score, relevance score, and security score; The accuracy score is determined based on the hallucination detection results; the relevance score is determined by calculating the cross-modal similarity between the user request, the original image visual features, and the VLM response using a cross-modal model; and the security score is determined based on whether the VLM response contains features of prompt injection attacks, unde-identified sensitive entities, or preset violation categories. The accuracy score, relevance score, and security score are weighted and summed according to preset weights to obtain the response quality score. Based on image-aware hashing, a historical response quality score baseline for similar images is maintained. Responses that deviate from the historical response quality score baseline, or have hallucination detection anomalies, abnormal security detection results, or negative user feedback, are identified as samples to be optimized.

9. The image intelligent preprocessing and privacy desensitization method for VLM inference according to claim 1, characterized in that, The user feedback data includes: Implicit feedback obtained through front-end interaction tracking points, business feedback obtained through downstream business metrics, and explicit ratings or negative reviews obtained through the user interface. Determining the image processing strategy to be optimized includes: associating and storing the hallucination type, preprocessing parameters, desensitized regions, response quality scores, and user feedback data in the samples to be optimized; For frequently occurring target categories, adjust the maximum long side parameter or the quality parameter of the format conversion for proportional non-scaled; Adjust the maximum long side, target format, quality parameters, or compression level parameters for preprocessing parameters that cause entity loss or blurred details; For samples with abnormal security scores, update the security restriction verification strategy and adjust the masking area or masking method based on samples with low relevance caused by masking.

10. An intelligent image preprocessing and privacy desensitization system for VLM inference, characterized in that, include: The request parsing and image recognition unit is used to receive VLM inference requests, recursively traverse the request body and match data URLs, and identify Base64 encoded image data. A security verification unit is used to perform format whitelist verification and security restriction verification on the image data; The image preprocessing unit is used to perform proportional scaling without magnification and format conversion or re-encoding on the verified image data; A privacy desensitization unit is used to perform region masking on the preprocessed image and update the VLM inference request in response to the VLM inference request being an assertion-type request. The request sending unit is used to send the updated VLM inference request to the upstream VLM service and obtain a VLM response. The metadata recording unit is used to record image processing metadata and VLM response associated metadata. In scenarios where verification fails or the request is terminated, the skip status and skip reason are recorded in advance. The hallucination detection unit is used to perform entity cross-validation on the VLM response to obtain hallucination detection results. The quality scoring unit is used to generate a response quality score based on hallucination detection results, semantic matching results, and security detection results. The feedback optimization unit is used to collect user feedback data and asynchronously perform hallucination detection, score calculation, feedback collection, and strategy update. Specifically, the processing strategies of the security verification unit, the image preprocessing unit, and the privacy desensitization unit are adjusted based on the image processing metadata, the hallucination detection results, the response quality score, and the user feedback data.