Image processing method and electronic device

CN122820872APending Publication Date: 2026-09-25LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611152526.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,目前这种图像处理方法需要将包含待处理对象完整视觉信息的图像数据上传到服务端,这很容易导致如精确轮廓、品牌Logo以及结构细节等核心视觉资产被第三方推断或提前曝光,尤其在涉及商业机密或个人定制物品等场景下,此类处理方式难以兼顾图像处理增强的数据安全性和高质量要求

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820872A_ABST
    Figure CN122820872A_ABST
Patent Text Reader

Abstract

The application discloses an image processing method and an electronic device, relates to the technical field of artificial intelligence, and specifically discloses the following technical scheme. A terminal device determines object feature data of a target object, the object feature data containing visual identity feature information of the target object. Based on the object feature data, reference feature data of the target object is generated, the reference feature data being used for representing target space attribute information of the target object in an image space and not containing visual feature information capable of restoring the visual identity of the target object. The reference feature data is sent to a server. Intermediate representation data generated by the server based on the reference feature data is received, the intermediate representation data being used for representing a distribution trend of visual elements in the image space. Thus, based on the object feature data and the intermediate representation data, a target image of the target object is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates primarily to the field of artificial intelligence technology, and more specifically to an image processing method and an electronic device. Background Technology

[0002] In scenarios such as image creation, industrial design, product appearance design, character design, game art setting, or brand visual asset design, users typically leverage the powerful computing resources of the server to perform high-quality rendering of the objects to be processed in the corresponding scene through the server-side model, in order to obtain display images with professional lighting and scene visual effects.

[0003] However, current image processing methods require uploading image data containing complete visual information of the object to be processed to the server. This can easily lead to core visual assets such as precise outlines, brand logos, and structural details being inferred or exposed in advance by third parties. Especially in scenarios involving trade secrets or personalized items, this type of processing method is difficult to balance the data security and high-quality requirements of image processing enhancement. Summary of the Invention

[0004] In view of the above problems, this application provides the following solution:

[0005] The first aspect of this application provides an image processing method, comprising:

[0006] Determine the object feature data of the target object; the object feature data includes the visual identity feature information of the target object;

[0007] Based on the object feature data, reference feature data of the target object is generated. The reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can restore the visual identity of the target object.

[0008] Send the reference feature data to the server;

[0009] Receive intermediate representation data generated by the server based on the reference feature data, wherein the intermediate representation data is used to characterize the distribution trend of visual elements in the image space;

[0010] Based on the object feature data and the intermediate representation data, a target image of the target object is generated.

[0011] A second aspect of this application provides an image processing method, comprising:

[0012] The terminal device receives reference feature data of a target object, wherein the reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can be used to reconstruct the visual identity of the target object;

[0013] Intermediate representation data is generated based on the reference feature data, and the intermediate representation data is used to characterize the distribution trend of visual elements in the image space;

[0014] The intermediate representation data is sent to the terminal device so that the terminal device generates a target image of the target object based on the object feature data of the target object and the intermediate representation data, wherein the object feature data includes the visual identity feature information of the target object.

[0015] A third aspect of this application provides an electronic device including at least one memory and at least one processor, wherein:

[0016] The memory is used to store the program of the intelligent agent;

[0017] The processor is used to run the agent's program to execute:

[0018] Determine the object feature data of the target object; the object feature data includes the visual identity feature information of the target object;

[0019] Based on the object feature data, reference feature data of the target object is generated. The reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can restore the visual identity of the target object.

[0020] Send the reference feature data to the server;

[0021] Receive intermediate representation data generated by the server based on the reference feature data, wherein the intermediate representation data is used to characterize the distribution trend of visual elements in the image space;

[0022] Based on the object feature data and the intermediate representation data, a target image of the target object is generated. Attached Figure Description

[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0024] Figure 1 This is a schematic diagram of the architecture of an image processing system proposed in this application;

[0025] Figure 2 This is a schematic flowchart of an image processing method proposed in Embodiment 1 of this application;

[0026] Figure 3 This is a schematic flowchart of an image processing method proposed in Embodiment 2 of this application;

[0027] Figure 4 This is a schematic diagram of visual potential field data in an image processing method proposed in an embodiment of this application;

[0028] Figure 5 This is a schematic flowchart of an image processing method proposed in Embodiment 3 of this application;

[0029] Figure 6 This is a schematic diagram of an image processing method for enhancing intermediate representation data, as proposed in an embodiment of this application.

[0030] Figure 7 This is a schematic diagram illustrating a boundary snapping activation process performed on intermediate representation data fed back from the server in an image processing method proposed in an embodiment of this application.

[0031] Figure 8 This is a schematic flowchart of an image processing method proposed in Embodiment 4 of this application;

[0032] Figure 9 This is a schematic flowchart of an image processing method proposed in Embodiment 5 of this application;

[0033] Figure 10 This is a schematic diagram of the structure of an image processing device according to Embodiment 1 of this application;

[0034] Figure 11 This is a schematic diagram of the structure of an image processing device according to Embodiment 2 of this application. Detailed Implementation

[0035] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments and is not intended to limit the application. The embodiments of this application are described below with reference to the accompanying drawings. It will be understood by those skilled in the art that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0036] The terms “first,” “second,” etc., used throughout this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0037] It is understood that before using the technical solutions disclosed in the embodiments of this application, users should be informed of the type, scope of use, and usage scenarios of the personal information (such as visual identity feature information of the target object) involved in this application, and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means. This application does not limit the implementation of prompt information and user authorization. In particular, this application protects user privacy through technical means (such as generating reference feature data that does not contain visual identity features), but users still need to know and agree to the entire process of the local device performing de-identification processing on the target object, the server generating intermediate representation data based on the de-identified data, and the local fusion generating the final image. The data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws and regulations and related provisions (such as corporate regulations). The server includes, but is not limited to, any one or a combination of the following: cloud architecture devices, local area networks and edge devices (terminal devices that provide computing power support within a local area network (LAN), edge computing nodes, edge gateways, home smart gateways, locally deployed private servers), and service nodes in end-to-end collaboration (other terminal devices that provide computing power support or data processing services in a local area network or local peer-to-peer network (P2P)).

[0038] Furthermore, the models involved in this application (such as the server-side generative model used to generate intermediate representation data, and the local model used for fusion rendering) can be general AI (Artificial Intelligence) models. AI models can employ, but are not limited to, Transformer or its architectural variants (such as encoder-only / decoder-only, encoder-decoder, or MoE (Mixture of Experts)) or other infrastructures. They learn the patterns of visual features and spatial layout by training on large amounts of diverse data, thereby enabling them to understand and generate image content. They typically have hundreds of millions to trillions of model parameters, capable of capturing complex relationships and patterns in images.

[0039] The AI ​​models can include, but are not limited to, generative models, diffusion models, generative adversarial networks (GANs), and visual Transformers (ViT). For example, one or more of the following: large language model (LLM), vision foundation model, and multimodal large model. In this application, the server-side model is specifically used to generate multi-channel visual potential fields (such as illumination potential field, density potential field, and background structure potential field) based on reference feature data, while the local model can be used to fuse and render object feature data with intermediate representation data.

[0040] Depending on actual needs, the model involved in this application can also be a large expert model fine-tuned from a general AI model based on actual business requirements, such as a proprietary model trained on a pre-trained model for a sample dataset of a specific product category (e.g., electronic products, handicrafts, industrial parts, etc.) or a specific rendering style (e.g., studio lighting effects, outdoor natural light, science fiction style, etc.). To meet the needs of edge deployment with limited computing resources and improve data security, this application can also compress the general AI model or large expert model through lightweight methods such as quantization, knowledge distillation, or pruning, and use the resulting lightweight model to implement the corresponding steps on the local side in the method of this application (e.g., reference feature data generation, local fusion rendering, etc.). In addition, the model involved in this application may also be a model expressed using certain rules or functions, such as physically based rendering equations or hand-designed visual field generation models, which can be determined according to actual needs. This application does not limit the type of model involved in the context.

[0041] Reference Figure 1 This is a schematic diagram of the architecture of an image processing system proposed in this application. The system may include a terminal device 110 and a server 120, etc. The terminal device 110 and the server 120 cooperate with each other, that is, through end-to-cloud collaboration, to implement the method proposed in the embodiments of this application. The form of the server 120 is as described above. Figure 1 The following description uses only the server-side 120 in server form as an example. In the image processing method proposed in this application described in the following embodiments, the server refers to the server-side component.

[0042] The terminal device 110 may have an image processing application installed or a web service accessible through a browser. The application or web page may provide an interactive interface through which the user inputs or uploads relevant operations and parameter configurations for the target object (such as electronic products, handicrafts or industrial parts to be displayed, etc., which need to be rendered and displayed; this application does not limit the type and form of the target object).

[0043] For example, terminal device 110 can receive original images or 3D models of target objects imported or captured by the user on the interactive interface, as well as rendering requirements set by the user. Terminal device 110 can generate reference feature data for the template object locally based on this input data and upload it to server 120. Since this reference feature data does not contain visual feature information that can reconstruct the visual identity of the target object, such as precise outer contour shape, brand logo, product model text, special hinge structure, interface opening position, screen ratio, surface texture pattern, etc., this avoids the risk of leakage of this visual feature information at the source during the intermediate representation data generation process by server 120, ensuring the information security of image processing.

[0044] Furthermore, since the server 120 returns intermediate representation data to the terminal device 110, this intermediate representation data can characterize the distribution trend of visual elements in the image space, thereby enabling the terminal device 110 to ultimately render and generate a high-quality target image based on the locally retained object feature data (which contains the visual identity feature information of the target object) and the intermediate representation data.

[0045] It should be understood that if the terminal device 110 has sufficient local computing resources, it can complete the entire image processing process on its own without the need for the server 120 to cooperate. This application does not impose any restrictions on this.

[0046] Optionally, the terminal device 110 in this application embodiment can be a mobile phone, tablet computer, smart wearable device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), desktop computer, or smart camera, etc., and this application does not limit its product form. In practical applications of this application, the terminal device 110 should have certain local image processing capabilities and network communication capabilities to complete the generation of reference feature data, the reception of intermediate representation data, and the final image fusion rendering and other processing.

[0047] In some embodiments, the terminal device 110 may include a radio frequency unit for wireless or wired communication with the server 120; a memory for storing a first computer program that implements the image processing method executed on the terminal side according to the present application, and data acquired or generated during the implementation of the image processing method, such as object feature data, reference feature data, intermediate representation data, and target image of the target object; a processor for executing the first computer program to implement the various steps of the image processing method proposed in the embodiments of the present application; an input unit such as a touch screen, mouse, keyboard, or stylus for receiving user operation instructions and parameter input; a display unit for displaying an interactive interface, preview image, and final target image; an image acquisition unit such as a camera for capturing the original image of the target object on-site; an audio circuit for voice interaction or prompts; an external interface for connecting external devices or data transmission; and at least one component such as a power supply. It is understood that these components are merely examples and do not constitute a limitation on the terminal or multifunctional device. The device may include more or fewer components, or a combination of certain components, or different components.

[0048] The server 120 may include a processor, a communication interface, a memory, and a bus for communication between these components. The memory in the server 120 may store a second computer program for the image processing method to be executed on the server side according to this application. This program is executed by the processor in the server 120 to implement the various steps of the image processing method on the server side. The processor may also schedule other units to implement corresponding functions, which will not be detailed here.

[0049] Based on the system architecture described above, the image processing method of this application embodiment will be described in detail below with reference to the accompanying drawings.

[0050] Reference Figure 2 This is a flowchart illustrating an image processing method proposed in Embodiment 1 of this application, applied to a terminal device. This terminal device can cooperate with a server to implement the method proposed in this application, such as... Figure 2 As shown, the image processing method proposed in this embodiment may include:

[0051] Step S21: Determine the object feature data of the target object.

[0052] In this application, the target object refers to any visual asset that the user wishes to enhance visually, and it can exist in various forms. For example, it can be an image such as a product photograph taken by the user, a character design drawing by a designer, or an electronic copy of a brand logo; a design file such as an industrial design sketch, a 3D model rendering, or a game art design draft; or a physical form in which the user obtains an image of the physical appearance of an object through scanning or photography and uses the obtained image as the target object for subsequent processing. This application does not limit the form in which the target object is represented.

[0053] Based on this, in the actual application of this application, the target object may include, but is not limited to, visual assets in the following application scenarios: such as the appearance of industrial products such as electronic products, automobiles or furniture; such as character IP designs such as original character concept art and three-view drawings; such as brand visual assets such as logos and brand identities; such as game or film art designs such as core character design drafts and scene concept art, which can be determined in combination with image processing requirements.

[0054] Object feature data refers to a set of data obtained from a target object that contains visual identity information of the target object. The specific content of object feature data can vary depending on the type of target object, application scenario, and user's personalized needs, but it always includes sensitive visual information that reflects the unique visual identity of the target object—that is, a set of key visual features that distinguish the target object from other objects. Based on these visual features, one can identify whether it is an unreleased product from a certain company or an IP character. For example, the appearance of an unreleased mobile phone, the brand logo on a product, the unique design of a character IP, or key outlines on industrial design drawings. Therefore, object feature data can include visual information that can uniquely or significantly identify the target object, thereby enabling the target object to be visually identifiable, distinguishable, or categorized.

[0055] The object feature data may include, but is not limited to, at least one of the following original spatial attribute information of the target object: geometric structure information, surface attribute information, spatial layout information, and identification area information. Geometric structure information reflects the physical form of the target object and is an important visual feature that distinguishes it from other objects. This includes at least one of the following: the target object's precise outline, boundary curves, interface opening locations, special hinge structures, and structural proportions. Surface attribute information reflects the surface texture of the target object and is an important parameter that needs to be matched in subsequent visual enhancement. This may include at least one of the target object's material type, texture features (such as surface texture patterns), and color distribution. Spatial layout information may include at least one of the target object's position in image space, orientation, and spatial relationship with surrounding elements. Identification area information may include at least one of the following: the area containing brand logos, text, or exclusive decorative elements on the target object (which may be referred to as the protected area).

[0056] Therefore, object characteristic data includes identity information indicating ownership, exclusive design information reflecting unique design, and core asset information carrying creative or commercial value. All types of information contained in object characteristic data are sensitive information requiring protection. If obtained by a third party, they could potentially identify the target object's identity, brand affiliation, or core design. Therefore, this application does not directly upload it to a server for processing but stores it locally. This non-external transmission method avoids the risk of the target object's core identity data being leaked to external networks from the source, thus improving the information security of the image processing process.

[0057] Based on the above analysis, the methods for determining object feature data can differ for different types of target objects. Optionally, this application can directly upload original images containing the target object through the interactive interface of the terminal device, such as product photos or design drawings created using design software, and the local processing system can use these as the source of object feature data. After the terminal device obtains the original image of the target object, it can automatically extract the object feature data of the target object from the original image using a locally deployed image processing model or AI model. For example, edge detection operators such as the Canny algorithm and the Sobel operator can be used to extract the outline of the target object from the original image; image segmentation operators based on thresholds or regions can be used to separate the target object from the background and obtain the occupied area / spatial occupancy relationship of the target object in the image space; and feature extraction networks of the model can be used to extract the texture features and color features of the target object. This application does not limit the implementation method of extracting object feature data through the model, and appropriate operators can be selected according to actual needs.

[0058] In another possible implementation, users can manually annotate various object feature data of the target object in the interactive interface. This includes drawing precise outlines on the product image using a mouse or stylus, selecting areas to be protected (such as logo areas, specially designed interface areas, etc.), or marking the visual focal point, such as the most eye-catching core area. Users can also determine the product's material type, such as metal, glass, or plastic, through drop-down menus or button selections. In practical applications, one or more of these methods can be used in combination, and the resulting object feature data can be displayed to the user for review, modification, or supplementary confirmation. This ensures that the core visual assets of the target object are protected from the source and not leaked to the server.

[0059] Step S22: Based on the object feature data, generate reference feature data for the target object. This reference feature data is used to characterize the target spatial attribute information of the target object in the image space, but does not contain visual feature information that can restore the visual identity of the target object.

[0060] In this embodiment, the reference feature data is generated based on object feature data and is used for subsequent processing on the server. To avoid leaking core visual assets during server processing, the object feature data determined above is first anonymized on the terminal device to retain the spatial attribute information of the target object in the image space (which can be called target spatial attribute information, to distinguish it from the original spatial attribute information of the target object in the image space). This provides the server with a spatial reference regarding the distribution of visual elements, such as "approximately where the target object is located, how large it is, and which direction it faces" in the image space, enabling the server to determine where the visual elements should be distributed. However, this target spatial attribute information itself does not contain visual details that can reconstruct the visual identity of the target object, such as the precise outline, logo, text, openings, and other visual identity features listed above. Therefore, the server cannot reconstruct or infer the true identity, brand, model, or unique structure of the target object based on this information.

[0061] Based on the above analysis and the object feature data described above, in some embodiments, the target spatial attribute information in the reference feature data may include: the approximate location of the target object in the image space (denoted as target location information), such as the center coordinates of the target object, the location range of the occupied area, etc.; the approximate size range of the target object in the image space (denoted as target size information), such as the width, height, area ratio, etc. of the target object; and the approximate shape of the target object in the image space (which may be denoted as target shape information), such as one or more of the target object's orientation trend, extension direction, etc.

[0062] It should be understood that the content of the object feature data that needs to be protected may differ depending on the image processing requirements of different users or scenarios. Therefore, the content of the generated reference feature data will also change accordingly, including, but not limited to, the target spatial attribute information listed in this embodiment. In practical applications, users can directly specify and upload the target spatial attribute information, or the local model can automatically generate reference feature data based on the object feature data. Alternatively, the generated reference feature data can be displayed to the user for review, revision, or supplementary confirmation to ensure that the reference feature data to be uploaded to the server does not contain sensitive information that needs protection, thus achieving flexibility in the reference feature data. This application does not limit the method of generating reference feature data.

[0063] As can be seen, the reference feature data generated in this application serves as an anonymous proxy for the target object. It retains the layout information required for the generated scene while completely masking the visual identity features, making it impossible for the server to reverse engineer the true appearance of the target object during subsequent processing.

[0064] Step S23: Send reference feature data to the server.

[0065] Step S24: Receive intermediate representation data generated by the server based on the reference feature data. This intermediate representation data is used to characterize the distribution trend of visual elements in the image space.

[0066] Following the above analysis, since the reference feature data does not contain the visual identity feature information of the target object, even if it is intercepted by a third party during transmission, the third party will not be able to obtain the true identity and core visual assets of the target object, thus ensuring the security of data transmission.

[0067] In some embodiments, the terminal device can send the reference feature data to the server via a wired or wireless communication network (such as the Internet, mobile communication network, or local area network). Optionally, the reference feature data can be encapsulated in a preset data format such as JSON (JavaScript Object Notation), XML (Extensible Markup Language), or binary format before being uploaded to the server.

[0068] Subsequently, the server can analyze the reference feature data using a pre-trained dedicated generative model or other models to generate intermediate representation data that characterizes the distribution trend of visual elements in the image space. It is evident that the intermediate representation data is not the final pixel image of the target object, but rather a data set describing how various visual elements should be distributed in the image space. This application does not restrict the method for generating the intermediate representation data.

[0069] Optionally, the intermediate representation data can describe the distribution trend of illumination in the image space, such as where the light comes from, how the illumination intensity changes in the image space, and how the warmth and coolness of colors are distributed in the image space; describe the distribution trend of textures or elements in the image space, such as how decorative particles are distributed in the image space, how textures are arranged in the image space; describe the distribution trend of background structures in the image space, such as the position and shape of the booth in the background of the target object, how the spatial hierarchy is divided, how the perspective direction extends, etc., one or more of these. This application does not limit the type of visual element distribution trend contained in the intermediate representation data. It can exist in the form of digital data and be transmitted from the server to the terminal device through a communication network as a transmittable data object.

[0070] Since the server only receives reference feature data (which does not contain visual identity feature information), and the intermediate representation data generated by the server also does not contain visual identity feature information of the target object, even if this data is intercepted by a third party, it is impossible to reconstruct the true identity or core visual assets of the target object. This application leverages the server's powerful computing capabilities and rich scene knowledge, coupled with the server's strong model generation capabilities, to quickly generate intermediate representation data with high-quality visual performance potential, without needing to know the identity of the target object, thus achieving efficient collaboration while protecting privacy.

[0071] Step S25: Generate a target image of the target object based on the object feature data and intermediate representation data.

[0072] After receiving the intermediate representation data returned by the server, the terminal device performs fusion processing by combining it with the object feature data held locally. That is, it uses the visual element distribution trend information in the intermediate representation data to provide a high-quality visually enhanced scene layout, and uses the visual identity feature information in the object feature data to make the scene layout accurately adapt to the target object. This ensures that the final generated visual elements such as light, shadow, texture, and background match the real identity and features of the target object, and then renders and generates a target image with high-quality visual enhancement effects that is accurately adapted to the target object.

[0073] The target image can be directly displayed on the terminal device's screen for user viewing, or saved as a digital image file and stored in local storage for later viewing; it can also be exported to specific applications or design tools for further editing or use; or it can be uploaded to other platforms for display or publication, etc. This application does not limit the output format of the target image.

[0074] In summary, in the image processing method proposed in this application, the visual identity feature information of the target object is always retained locally. The reference feature data uploaded to the server only contains the desensitized target spatial attribute data and does not contain sensitive information that can restore the visual identity of the target object, thus fundamentally eliminating the risk of visual identity leakage. The server is responsible for generating large-scale scene layouts (intermediate representation data that is not pixel information), while the terminal device is responsible for preserving visual identity features and generating the final target image locally. This approach protects the privacy and security of the target object while leveraging the server's efficient rendering capabilities and powerful computing power, giving full play to the advantages of both parties, reducing the local computing burden, and ensuring that the final generated target image achieves professional-level visual enhancement quality, possessing both realism and artistic appeal. Therefore, the edge-cloud collaborative processing method adopted in this application avoids both the insufficient computing power problem faced when relying on local models and the privacy leakage risk faced when directly uploading original images, achieving an effective balance between privacy protection, computing efficiency, and image generation quality.

[0075] Based on the above analysis, the implementation process of this application's method is illustrated using the scenario of generating promotional images for an unreleased product as an example. Assume a user is generating promotional images for an unreleased smartphone from a certain brand. The user imports the official render image of the phone into the local processing system of the terminal device to determine the phone's object feature data, such as the phone's precise outline, camera module position, brand logo area, and material partitioning between the metal frame and glass back panel. Subsequently, target spatial attribute information such as the phone's position area, size range, and orientation trend in the image can be extracted and uploaded to the server as at least a portion of the reference feature data. Sensitive information such as the phone's actual appearance details, logo position, and precise outline are not uploaded, avoiding the risk of leakage during upload or server processing.

[0076] After receiving the intermediate representation data from the server, the terminal device can precisely adapt the intermediate representation data based on the phone's object feature data. This includes precise alignment of edge lighting with the phone's contours, matching light reflection to metal / glass materials, and avoiding logo areas with particle decorations. Based on the adapted intermediate representation data, a product display image with a professional studio-quality feel is generated. Throughout this process, the phone's actual appearance design is never uploaded to the server, eliminating any risk of leakage. Applying this method in other application scenarios can achieve visual effects far exceeding the performance of local devices while protecting the privacy and security of the target object. Detailed examples are not provided here.

[0077] Reference Figure 3This is a flowchart illustrating an image processing method proposed in Embodiment 2 of this application. This embodiment describes a possible implementation method of how a terminal device generates reference feature data of a target object based on object feature data, such as... Figure 3 As shown, the implementation method may include:

[0078] Step S31: Extract the original spatial attribute information of the target object in the image space from the object feature data.

[0079] Based on the above description of target spatial attribute information, the original spatial attribute information extracted from object feature data refers to the original data related to the target object's position, size, shape, and other attributes in image space. It contains fine spatial details of the target object, and the level of detail (such as the density of contour point sequences and the degree of preservation of curve details) may enable third parties to identify the visual identity of the target object.

[0080] Optionally, the original spatial attribute information may include, but is not limited to: the precise contour point (pixel-level) sequence of the target object, the original bounding rectangle parameters, the projection contour of the depth map onto the image plane, the projection range of the 3D model onto the image plane, or the object region boundary in the semantic segmentation mask, etc. This information comes directly from the original image or 3D scan data of the target object, has high accuracy, and contains a large number of details that can be used to identify the target object. It needs to undergo subsequent de-identification and dimensionality reduction processing before being uploaded to the server.

[0081] For example, if the target object is a smartphone, its original spatial attribute information may include: the precise outline curve of the screen bezel, the relative position of the camera module, the coordinates of the button openings, and the pixel area of ​​the brand logo. Although this information falls under the category of spatial attributes, the precise geometric features it contains are sufficient to uniquely identify the phone's model and even its specific batch.

[0082] In some embodiments, one or more combinations of contour extraction, extrinsic parameter extraction, depth information extraction, and user-annotated extraction methods can be used to extract the original spatial attribute information. Specifically, in the contour extraction process, an edge detection operator can be used to extract a precise sequence of contour points of the target object, which includes fine geometric details of the target object's boundary, such as subtle undulations of curves, radian variations at specific angles, and the concave-convex structure of local areas. Furthermore, the relative positional relationships between different structural components can be determined.

[0083] In the implementation of the circumscribed parameter extraction method, the circumscribed rectangle parameters of the target object in the image space, such as the coordinates of the upper left corner, width, and height, can be calculated through the target detection model or mathematical operations. Alternatively, the circumscribed ellipse parameters of the target object in the image space can be calculated, such as the center coordinates, major and minor axis lengths, rotation angle, and principal axis direction angle.

[0084] When the object feature data includes depth maps or point cloud data, a depth information extraction method can be used. This involves projecting the depth map or point cloud data onto the image plane to obtain the projected outline or occupied area (projection range) of the target object in the two-dimensional image space. Furthermore, the terminal device can also respond to user annotation operations on the interactive interface, acquiring the original spatial attribute information such as the outline, bounding box, or key point positions of the user-annotated target object. This application does not limit the implementation method of step S31.

[0085] Step S32: Remove visual feature information that can restore visual identity features from the original spatial attribute information and perform dimensionality reduction processing on the information to obtain the target spatial attribute information of the target object.

[0086] Step S33: Determine the target spatial attribute information as at least a portion of the reference feature data to be transmitted.

[0087] To protect privacy and data security, this application can remove or blur sensitive visual features in the original spatial attribute information that can restore the visual identity of the target object during the desensitization process. That is, it can identify and remove the geometric details that are strongly related to the identity of the target object as described above. This application does not limit the desensitization process.

[0088] Optionally, the precise outline can be smoothed or simplified to eliminate identifiable features such as sharp corners and unique curvatures; coordinate information of iconic areas such as brand logos, model text, and special openings can be removed; specific dimensions such as screen ratio and button spacing can be normalized or blurred; and fine component segmentation results can be merged into coarse-grained overall occupying areas, thereby completely removing the visual identity features in the original spatial attribute information that can be used to infer the specific model or brand of the target object.

[0089] Since the spatial attribute information after the above-mentioned desensitization processing is usually still high-dimensional data, this application can further perform dimensionality reduction processing to reduce its spatial complexity. Spatial complexity refers to the amount of data or information entropy required to describe the spatial attributes of the target object (which reflects the level of detail). It can be measured by at least one of the following indicators: the number of contour points, the fitting order, the amount of data stored, or the level of detail in the geometric description. Lower spatial complexity indicates a coarser and more generalized description of the spatial attributes of the target object; higher spatial complexity indicates a more detailed and closer description of the spatial attributes of the target object to its original form.

[0090] Optionally, the dimensionality reduction process described above can employ one or more combinations of the following methods, but not limited to: resolution reduction, which involves downsampling the original high-resolution contour image (e.g., 1024×1024) into a low-resolution placeholder image (e.g., 64×64), retaining only the macroscopic placeholder trend; parametric simplification, which involves converting the pixel-level contour into a few geometric parameters, such as center coordinates, the width and height of the bounding rectangle, and the principal axis direction angle, discarding all non-parametric details; sparse representation, which involves retaining only the key anchor points of the target object in the image space, such as the four vertices and centroid of the target object, discarding all other intermediate points; and frequency domain truncation, which involves performing a Fourier transform on the original spatial attribute information, retaining only low-frequency components, and filtering out high-frequency details.

[0091] Through dimensionality reduction, the spatial complexity of the obtained target spatial attribute information is significantly lower than that of the original spatial attribute information. For example, the original spatial attribute information may require thousands of bytes to describe a precise contour, while the dimensionality-reduced target spatial attribute information may only require tens of bytes of coordinate parameters to characterize it. This application does not restrict the implementation method of desensitization and dimensionality reduction of the original spatial attribute information.

[0092] In some embodiments, the desensitization process (i.e., removing visual feature information that can restore visual identity features) and dimensionality reduction process in step S32 can be implemented simultaneously. The implementation of desensitization and dimensionality reduction processes for the original contour point sequence is described below as an example. The processing methods for other types of original spatial attribute information are similar, and this application will not provide detailed examples for each.

[0093] Optionally, the original contour point sequence can be processed using a bounding rectangle fitting method. For example, the minimum bounding rectangle algorithm can be used to calculate the minimum bounding rectangle of the original contour, retaining only the center coordinates, width, and height parameters of this rectangle. In this way, the spatial attributes of the target object in image space are represented by four parameters: center x-coordinate, center y-coordinate, width, and height, making its spatial complexity much lower than the original contour point sequence (which may contain hundreds to thousands of coordinate points). It is evident that the bounding rectangle only describes the approximate location and size range of the target object, losing the shape details of the contour; therefore, third parties cannot reconstruct the visual identity of the target object based on this.

[0094] In another possible implementation, this application can also employ ellipse fitting, using the least squares method to fit the original contour into an ellipse, retaining only the center coordinates, major axis length, minor axis length, and rotation angle parameters of the ellipse. Its spatial complexity is far lower than that of the original contour point sequence. It is evident that ellipse fitting preserves the general orientation trend while removing the fine undulations and local features of the contour.

[0095] In another possible implementation, this application can also employ polygon simplification, such as the Ramer-Douglas-Peucker algorithm, to simplify the original contour point sequence, retaining major inflection points and removing minor undulations. By setting a simplification threshold, the number of vertices in the simplified polygon can be controlled. The simplified polygon has far fewer vertices than the original contour points, significantly reducing spatial complexity. Because the fine details of the contour are removed during the simplification process, third parties cannot reconstruct the precise shape of the target object based on this.

[0096] Furthermore, the desensitization and dimensionality reduction of the original contour point sequence can be achieved using a binary mask generation method. This involves generating a binary image where the area occupied by the target object is represented by a first pixel value (e.g., 255, white), and the remaining areas are represented by a second pixel value (e.g., 0, black). This binary mask image only describes the area occupied by the target object in the image space and does not contain any internal details of the target object, such as color, texture, or brand logo. Moreover, the data size of the binary mask depends on the image resolution, but it only describes binary information of occupancy / non-occupancy, resulting in a spatial complexity far lower than the fine geometric information contained in the original contour point sequence.

[0097] Optionally, this application can also employ downsampling to sample the original contour point sequence at equal intervals, reducing the number of contour points. For example, if the original contour contains 1000 points, it can be reduced to 100 points by sampling one point for every 10 points. The downsampled contour point sequence retains the general shape of the target object but loses fine local details, significantly reducing spatial complexity. Alternatively, Gaussian blurring can be used to blur the original contour, making its fine structure indistinguishable. Gaussian blurring smooths out subtle undulations and local irregularities of the contour by convolving the contour point coordinates with a Gaussian kernel, retaining only the overall direction of the contour. The implementation process is not detailed in this application.

[0098] It should be noted that the desensitization and dimensionality reduction methods listed above can be used individually or in combination. For example, the original contour point sequence can be downsampled first, and then the bounding rectangle of the downsampled contour can be calculated—this is desensitization. Further dimensionality reduction can also be performed. Regardless of the method or combination used, the goal is to reduce the spatial complexity of the original spatial attribute information while removing sensitive details that could reconstruct the visual identity of the target object, resulting in target spatial attribute information containing only approximate location, size, and orientation trends. This significantly reduces the data volume, facilitating efficient transmission.

[0099] In this application, reference feature data refers to the data set sent by the terminal device to the server for the server to generate intermediate representation data. In some embodiments, the target spatial attribute information obtained by the above processing can be directly determined as reference feature data and uploaded to the server. In other embodiments, the processed target spatial attribute information is only a part of the reference feature data, and the reference feature data may also include other auxiliary information (which is non-sensitive information), such as timestamps, session identifiers, non-sensitive background information, or rendering requirement information for the target object. The rendering requirement information may include, but is not limited to, at least one of the following: target style, lighting direction, color atmosphere, background type, decoration density, display angle, screen ratio, and generation resolution.

[0100] In summary, this application employs two-stage processing—removing visual identity features and dimensionality reduction—to ensure that even if the reference feature data is intercepted by a third party, the true appearance or model of the target object cannot be recovered. The dimensionality-reduced target spatial attribute information is extremely small, which not only reduces network bandwidth consumption and latency but also decreases the amount of data processed by the server, improving the overall efficiency of end-to-cloud collaboration. Furthermore, despite the significant reduction in information volume, the target spatial attribute information still retains the approximate location, size range, and orientation trend of the target object in image space. This information is sufficient for the server to determine where visual elements should be distributed, providing the necessary spatial reference for the server to generate high-quality intermediate representation data without affecting subsequent rendering effects.

[0101] Furthermore, this application provides the aforementioned various de-identification and dimensionality reduction processing methods. Users or system developers can flexibly choose appropriate processing methods or combinations according to different requirements for privacy protection strength and spatial reference accuracy in application scenarios, exhibiting good adaptability and scalability. It should be noted that the processing methods described above can be implemented by users using relevant software / tools, or can be automatically executed by terminal devices based on preset rules or processing models; this application does not impose any restrictions on this.

[0102] In some embodiments, the intermediate representation data generated by the server may include visual potential field data, which is a multi-channel data set used to characterize the distribution trend of visual elements in image space. It is clear that the visual potential field data in this application is not the final pixel image, but a descriptive intermediate representation data that describes where light should come from, how decorative elements should be distributed, and how the background structure should be organized in image space, but does not contain the specific pixel content of the target object. Based on this, the visual potential field data includes at least one of the following sub-field data: illumination potential field data, density potential field data, and background structure potential field data. Each sub-field data describes the distribution trend of visual elements in different dimensions of image space, is independently encoded, and aligned under a unified image space coordinate system. Each sub-field data can be stored in the form of a multi-dimensional array (such as a tensor of H×W×C), where H and W represent the number of grid cells in the height and width directions of the image space, respectively, corresponding to the target output resolution. Their values ​​are typically greater than the number of grid cells in the corresponding direction of the target object in the image space for subsequent fine-tuning. C represents the number of channels in the sub-field data, and this application does not limit its value.

[0103] In some embodiments, illumination potential field data can be used to characterize the distribution of light intensity, chromaticity, or ray direction parameters in the image space. It describes what kind of illumination each location in the image space should receive, including the direction from which the light comes, the intensity of the light, and the color bias, providing a lighting reference for subsequent local rendering to ensure that the final generated target image has professional-grade lighting effects. For this purpose, the illumination potential field data may include at least one channel data from the ray direction channel, the intensity channel, and the chromaticity channel to determine the number of channels C.

[0104] Among them, the light direction channel data is two-dimensional vector data, which describes the direction of the main light received at that location, such as the vector direction angle within the range of 0°-360°. It can be seen that this channel data can record the incident direction of the main light at each grid cell, represented by a two-dimensional vector or a normalized direction vector. It determines the projection direction of light and shadow in the scene and directly affects the brightness distribution and shadow shape of the target object.

[0105] The light intensity channel data is a scalar data type that describes the light intensity at a given location and can be a normalized value within the range of 0-1. In this application, this channel data can also be called material texture channel data, which records the surface material response characteristics at each mesh cell, such as diffuse reflection coefficient and specular reflection intensity, to simulate the optical performance of different materials such as metal, plastic, or fabric under illumination, so as to ensure that the final generated target image presents the real material of the target object.

[0106] Chromaticity channel data is a type of three-dimensional vector data, such as RGB channel data, which describes the lighting color at that location and can be a value within the range of 0-255 or 0-1. Therefore, this channel data can also be called color gradient channel data, recording the ambient light color and color temperature gradient at each grid cell. For example, it can transition from warm tones (such as 3000K) to cool tones (such as 6500K) from left to right, or simulate the orange-blue gradient of the sky at dusk, which can be determined according to the actual situation.

[0107] As can be seen, the illumination potential field data proposed in this application can adopt an H×W×3 tensor format, with the three channels corresponding to the light direction, light intensity, and chromaticity, respectively. Optionally, the illumination potential field data can adopt an H×W×6 tensor format, where the first two channels are two-dimensional vectors of the light direction, such as the x-direction and y-direction components, the middle three channels are the RGB values ​​of the chromaticity, and the last channel is the light intensity value. The three channels correspond to the light direction, light intensity, and chromaticity, respectively. Different channel configurations can be selected according to actual needs, and this application does not impose any restrictions on this.

[0108] After receiving the reference feature data of the target object and the rendering requirement information, the server-side model can generate illumination potential field data based on the target spatial attribute information and the rendering requirement information. Combining the lighting direction and color atmosphere parameters in the rendering requirements, a global illumination distribution is generated through a spatial attenuation function. For example, mesh cells closer to the target object's occupied area receive stronger illumination response, while those farther away gradually become ambient light. This application embodiment does not limit the method for generating illumination potential field data; please refer to the description of the corresponding part of the server-side execution embodiment below. It should be noted that the spatial distribution range of the illumination potential field is referenced to the spatial occupied area of ​​the target object, but the calculation of its internal pixel values ​​does not depend on any identity feature information of the target object.

[0109] For example, if a user wants to generate a display image for a smartwatch, the rendering requirements specify the lighting direction as "side backlighting" and the color atmosphere parameter as "cool tone." In the lighting potential field data generated by the server, the light intensity channel value of the upper left region of the image (which is the direction of the light source) is high, such as 0.9, and the chroma channel is cool, such as R=180, G=220, B=255, etc.; the light intensity channel value of the lower right region of the image (which is the direction of the backlight) is low, such as 0.2, and the chroma channel is dark, such as R=60, G=80, B=100, etc.; the light direction channel of the area where the target object is located points to the upper left, indicating that the light received in this area comes from the upper left. Thus, when the local rendering is performed based on this lighting potential field data, the top and left sides of the smartwatch will have highlights, and the bottom and right sides will have shadows, presenting an overall cool-toned side backlighting effect. Similarly, if the user specifies "warm light at 45 degrees to the left", the illumination potential field data will generate a high-brightness warm light vector on the left side of the occupied area and a gradual cool shadow transition on the right side, while assigning a moderate diffuse reflection coefficient to the occupied area in the material texture channel.

[0110] The density potential field data in the intermediate representation can be used to characterize the distribution density of particles, textures, or lines in the image space. It describes how many decorative elements, such as floating light spots, starlight, particles, texture patterns, or lines, should be distributed at each location in the image space, providing a reference for the layout of decorative elements in subsequent local rendering. It can be seen that the density potential field data can include at least one channel data from particle density channels, texture density channels, and line density channels. The channel configuration can be selected according to actual needs, preferably using a three-dimensional array structure of H×W×3.

[0111] The particle density channel records the probability or density of decorative particles (such as light spots, dust, or snowflakes) in each grid cell, expressed as a scalar value within the range of 0-1. High-density areas will render dense particle effects, while low-density areas will be cleaner. The texture density channel records the repetition frequency or coverage intensity of background textures (such as wood grain, marble, or fabric textures) in each grid cell, expressed as a scalar value within the range of 0-1. This channel can control the smoothness of the background surface. The line density channel records the distribution density of structural lines (such as futuristic light trails, architectural outlines, or decorative borders) in each grid cell, expressed as a scalar value within the range of 0-1. This channel can be used to create rhythm and visual guidance in the image.

[0112] Based on this, when the server-side model generates density potential field data according to the target spatial attribute information and rendering requirement information in the reference feature data, it can construct a radial density distribution function centered on the target object's occupied area. In the inner region immediately adjacent to the target object's occupied area, i.e., the grid area occupied by the target object itself, the density value is set to the lowest, approaching zero, to ensure that background decorations do not intrude into the target object's body. In the middle layer region surrounding the target object's occupied area, the density value gradually increases, forming a natural decorative transition. In the edge region far from the target object's occupied area, the density value reaches its highest, used to fill the blank areas of the image. This density distribution pattern ensures that the visual center of gravity is always concentrated around the target object, but it is not limited to this method of generating density potential field data.

[0113] For example, suppose a user wants to generate a display image for a pair of sneakers (the target object). The rendering requirements specify a cyberpunk style and moderate density. In the density potential field data generated by the server, the density value of the area where the sneakers are located is 0, indicating that no decorative elements cover the shoes. The density value in the 0-30 pixel range around the sneakers gradually changes from 0 to 0.5 to form a transition zone between the sneakers and the background area. The density value of the background area is 0.5-0.7 to represent a stable distribution of decorative elements. High-density areas are concentrated at the four corners and edges of the image, forming decorative rings around the subject. Thus, when the local rendering is performed based on this density potential field data, cyberpunk-style neon lines, grid textures, and floating data points will appear around the sneakers and in the background area, while the sneakers themselves remain clearly visible. Similarly, if the user specifies a sci-fi style and medium decoration density in the rendering requirements information, the density potential field data will generate dense sci-fi light trails (high value of the line density channel) around the target object's occupied area, while a small number of light spot particles (medium value of the particle density channel) will be scattered at the edge of the image, while the target object's occupied area will remain at zero density.

[0114] The background structure potential field data in the intermediate representation data is used to characterize the distribution of spatial hierarchy, geometric structure or background topology in the image space. It describes the spatial structure of the background region in the image space, such as which regions belong to the foreground, which belong to the middle ground, which belong to the background, how the background lines should extend, and the perspective relationship in the space, etc., providing a spatial reference for the background layout for subsequent local rendering.

[0115] As can be seen, background structural potential field data can include at least one of the following channels: spatial hierarchy channel, geometric type channel, and structural direction channel. For example, using an H×W×3 tensor format, the first spatial hierarchy channel is a scalar, which can record the depth level or depth-of-field information of each grid cell. For instance, foreground layer, subject layer, background layer, and infinity layer are used to control the sense of depth and the relationship between sharpness and blur in the image, determining which areas should be clear and which should be blurred (such as depth-of-field effects). This can be controlled by values ​​within the range of 0 (foreground) to 1 (background). The latter two channels are two-dimensional vectors representing the structural direction (e.g., x- and y-direction components). The structural direction channel describes the extension direction of the structural lines at that location and can be represented by a vector direction angle within the range of 0°-360°. The geometric type channel describes the geometric construction type at that location, such as a stand, a plane, a curved surface, or a cylinder. Therefore, the geometric construction channel data can record the background geometric structure information at each grid cell, such as the tilt angle of the stand, the perspective direction of the wall, and the grid orientation of the ground. It can be represented by a two-dimensional vector (dx, dy) to indicate the offset or orientation of each grid cell relative to the overall spatial structure, used to construct the perspective consistency of the background. Of course, the background structural potential field data can also adopt an H×W×2 tensor format, containing only the spatial hierarchy channel and the structural direction channel. Different channel configuration methods can be selected according to actual needs, and this application does not impose any restrictions on this.

[0116] Based on this, during the process of generating background structural potential field data using target spatial attribute information and rendering requirement information from the reference feature data, the server-side model can automatically plan the position and orientation of the booth or supporting structure according to the target object's occupied area. The booth is usually located directly below the target object's occupied area, and its size matches the outer boundary of the occupied area. The spatial hierarchy channel decreases in depth layer by layer outward from the target object's occupied area, forming a natural depth-of-field transition. The geometric construction channel generates the corresponding perspective mesh or structural vector field based on the background type in the rendering requirement information, such as a solid color background, an indoor scene, or an outdoor landscape.

[0117] For example, suppose a user wants to generate a display image for a headphone product (i.e., the target object), and the rendering requirements specify a head-up viewing angle and a booth background type. In the background structure potential field data generated by the server, the area where the headphones are located is the foreground layer with a level value of 0; the area below and behind the headphones is the midground layer with a level value ranging from 0.4 to 0.6, forming a rectangular booth area; the area behind the booth is the background layer with a level value ranging from 0.8 to 1.0, resulting in a smooth transition of spatial layers; the structural direction channel shows that the background lines extend horizontally outward from the headphone position, forming a stable horizontal composition. When the local device renders the image based on this background structure potential field data, the headphones will be presented on a horizontal booth, with the background lines extending horizontally, resulting in a stable and professional overall composition. Similarly, if the user specifies an indoor booth and a top-down angle, the background structural potential field data will generate an inclined booth plane below the target object's occupied area. At this time, the booth area vector in the geometric construction channel points to the top-down direction, and the booth is set as the middle layer, the ground as the bottom layer, and the wall as the upper layer in the spatial hierarchy channel, forming a clear indoor spatial hierarchy.

[0118] Therefore, it is evident that the illumination potential field data, density potential field data, and background structure potential field data are independent yet synergistic, collectively constituting a complete visual potential field data set. This data comprehensively describes the distribution trends of visual elements in the image space from three dimensions: illumination, decoration, and structure. Upon receiving this visual potential field data, the local device can combine it with the visual identity features of the target object to modulate the three sub-field data separately, ultimately rendering a target image with high-quality visual effects.

[0119] Reference Figure 4 The diagram shows a visual potential field data, where the white sphere represents the reference feature data of the target object. It represents the macroscopic occupancy of the target object in the image space. It appears as a concave area in the diagram, symbolizing that the space in the visual potential field has been physically occupied by the target object, and background elements need to avoid this area. Figure 4 The colored surface in the image represents the multi-channel visual potential field. It is a colored undulating surface overlaid on the grid plane, representing the visual potential field data generated by the server. Each height (Z-axis) and color change of the surface encodes specific visual rendering parameters, such as light intensity, texture density, background depth, and other sub-field data. It should be noted that the three sub-field data are aligned in a unified image space coordinate system. That is, the same spatial coordinate corresponds to the same pixel position in image space in different sub-field data, which allows the local end to perform superposition, modulation, and fusion of the three sub-field data under the same spatial reference system during subsequent processing.

[0120] Among them, such as Figure 4As shown, in the density potential field data, the value at the depression where the sphere is located approaches zero. This indicates that the server-side model should not fill this area with particles, textures, or lines, thus visually hollowing out the area where the target object is located, ensuring sufficient white space or a clean background around it, and preventing background elements from obscuring the subject. In the illumination potential field data or background structure potential field data, the raised peaks immediately adjacent to the edge of the sphere represent visual guidance. In the illumination potential field data, the peaks may represent highlights hitting the edge of the object, forming rim lights and enhancing the sense of depth. In the background structure potential field data, the peaks may represent the edges of the display stand, halos, or perspective vanishing points in the background used to highlight the subject, forming the edge of a gravity well that locks visual attention to the area around the central sphere.

[0121] Therefore, the server-side model receives a simple geometric placeholder of the target object (such as...). Figure 4 The white sphere shown doesn't need to be known in detail; it could be a mobile phone, computer, or other product. By analyzing the spatial disturbances (such as depressions) generated by the sphere, the potential energy of the surrounding environment (surface undulations) can be calculated to generate the optimal lighting layout and background topology. This allows the terminal device to place local real objects, such as mobile phones with complex textures and colors, into this pre-calculated visual potential field, thus completing the final physically-based rendering.

[0122] In summary, this application decomposes complex visual scenes into three semantically defined subfields, facilitating targeted generation and adjustment of the server-side model and enabling local modifications by the terminal device in subsequent steps, achieving refined visual enhancement control. The visual potential field data is stored in the form of a multidimensional array (tensor), with a compact and standardized data format, facilitating transmission over the network and local processing. Compared to transmitting a complete pixel image, transmitting visual potential field data requires less data and is more efficient. The three subfields are aligned in a unified image space coordinate system, ensuring the consistency of the spatial distribution of visual elements in each dimension and preventing misalignment of illumination, density, and background structure. Furthermore, none of the three subfields contain any visual identity features of the target object; even if all of them were captured, the true appearance of the target object could not be reconstructed.

[0123] The visual potential field data includes at least one subfield data. Users can choose to use all three subfield data, or only one or two, depending on the specific application scenario. For example, in scenarios requiring only illumination adjustment, only the illumination potential field data can be used; in scenarios requiring comprehensive enhancement, all three subfield data can be used simultaneously. This flexible configuration allows the solution to adapt to different application needs.

[0124] Based on the above analysis, referring to Figure 5The diagram shown is a flowchart of an image processing method proposed in Embodiment 3 of this application. This embodiment describes a possible implementation method for generating a target image of a target object based on object feature data and intermediate representation data. Figure 5 As shown, the implementation method may include:

[0125] Step S51: In response to the interactive operation of the target object in the image space, the intermediate representation data is adjusted to obtain the adjusted intermediate representation data.

[0126] Based on the above analysis, the intermediate representation data returned by the server, namely the visual potential field data, roughly matches the position and shape of the target object. Since the server does not obtain the object feature data of the target object, in order to obtain a high-quality target image, it is necessary to fine-tune it locally using real object feature data. For example, users can make preliminary adjustments to the visual potential field data returned by the server according to their personal aesthetic preferences or specific display needs, so that its spatial layout better meets the user's expectations. Then, the system will perform precise adaptation based on the visual identity feature information of the target object.

[0127] In one possible implementation, users can adjust the layout of visual elements indicated by intermediate data through the interactive interface of the terminal device. This can be achieved through direct operations such as dragging, scaling, and rotating to adjust at least one of the target object's position, size, and orientation in the image space. These spatial transformation-type interactive operations can be implemented using input components such as mouse dragging, touch gestures, or a keyboard. For example, a user can drag the target object from the center of the screen to the lower right corner, enlarge it to occupy two-thirds of the screen area, or rotate it 30 degrees to present a better viewing angle.

[0128] Each interactive operation can trigger a real-time update of the intermediate representation data and display a preview image of the target object generated based on the updated intermediate representation data. In response to the above interactive operation, the corresponding subfield data in the intermediate representation data undergoes spatial remapping, so that the illumination potential field data, density potential field data and / or background structure potential field data are adjusted synchronously with the spatial transformation of the target object. That is, based on the new occupancy area of ​​the target object in the image space, the spatial distribution of the illumination potential field, density potential field and background structure potential field is recalculated to match the scene layout with the current occupancy of the target object.

[0129] For example, when a user moves a target object from its current position to a new position in the image space by dragging with a mouse or swiping with a touch, the terminal device responds to the drag operation by calculating the displacement vector of the target object. Based on this displacement vector, it translates the subfield data in the intermediate representation data, essentially resampling the pixel values ​​in each subfield according to the displacement vector. After the translation, the illumination distribution in the illumination potential field data, the decoration density distribution in the density potential field data, and the spatial hierarchy distribution in the background structure potential field data all move to the new position along with the target object, maintaining their relative spatial relationship.

[0130] Users can zoom in or out of the target object in image space by dragging its boundary control points or using a pinch gesture. In response to this zooming operation, the terminal device calculates the scaling ratio of the target object and performs a scaling transformation on the subfield data in the intermediate representation data according to this ratio. This involves resampling and interpolating the subfield data according to the scaling ratio. After scaling, the illumination transition range in the illumination potential field data, the distribution range of decorative elements in the density potential field data, and the spatial hierarchy range in the background structure potential field data all adjust synchronously with the change in the size of the target object.

[0131] Users adjust the orientation of a target object in image space through rotation operations, such as dragging a rotation handle or inputting a rotation angle. The terminal device responds to this rotation operation by calculating the rotation angle of the target object and performing a rotation transformation on the subfield data in the intermediate representation data based on this rotation angle; that is, resampling the subfield data according to the rotation angle. After rotation, the light direction distribution in the illumination potential field data (such as the relative angle between the main light source direction and the target object) and the structural direction in the background structural potential field data (such as the extension direction of background lines) are adjusted synchronously with the change in the orientation of the target object.

[0132] It should be noted that the implementation methods of the above-mentioned spatial transformation interactive operation include, but are not limited to, using mouse and keyboard or touch control. Other input devices that are connected to the terminal device, such as smart gloves or control handles, can also be used to realize the spatial transformation interactive operation of the target object. The implementation process will not be described in detail in this application.

[0133] In another possible implementation, the above-mentioned interactive operation can also be a cue word-based interactive operation. In this case, the user can input adjustment cue words for the intermediate representation data through natural language, so that the local model can make local or global adjustments to the corresponding subfields of the intermediate representation data based on the adjustment cue words. Here, the natural language can be in the form of text or speech, etc., and the adjustment cue words can include at least one of the following: illumination adjustment cue words, density adjustment cue words, and background structure adjustment cue words.

[0134] For example, if a user inputs adjustment prompts such as "adjust the main light source direction to 45 degrees to the left," "enhance the warm color atmosphere," or "increase the softness of the background light," the local model parses these prompts, identifies the type of subfield data the user wants to adjust (in this case, illumination potential field data), and the specific adjustment parameters, such as light source direction, color temperature, and softness. Based on this parsing result, the model updates the parameter values ​​of the corresponding channels in the illumination potential field data. Similarly, if a user inputs adjustment prompts such as "increase the particle density in the background," "reduce the number of decorative lines," or "add floating light spots around the subject," the local agent or model identifies the type of subfield data the user wants to adjust (i.e., density potential field data) and the specific adjustment parameters, such as density value and element type. Based on this parsing result, the model updates the density value or element type identifier at the corresponding position in the density potential field data. User input prompts such as "change the background to a blurred effect," "increase the sense of depth in the booth," or "make the background lines converge towards the center." The local agent or model identifies the subfield data type (i.e., background structural potential field data) that the user wishes to adjust, as well as the specific adjustment parameters, such as spatial hierarchy distribution and structural direction. Based on this analysis, the agent updates the corresponding spatial hierarchy value or structural direction vector in the background structural potential field data. This application does not restrict the entity that performs the initial adjustment of the intermediate representation data.

[0135] After the above interactive operations, the adjusted intermediate representation data can reflect the user's initial intention regarding the scene layout and atmosphere, but it still does not contain the true visual identity features of the target object. Further fine-tuning can be performed according to the steps described below. At this point, the adjusted intermediate representation data has the same data structure as the original intermediate representation data, namely the same number of subfield data channels and the same image space coordinate system. However, the numerical distribution of each subfield data shows changes corresponding to the user's operation. For example, if the user shifts the target object 50 pixels to the right, in the adjusted intermediate representation data, the illumination potential field data, density potential field data, and background structure potential field data are all shifted 50 pixels to the right. The values ​​of each subfield data at each pixel position are redistributed, but the relative spatial relationships between the values ​​remain unchanged.

[0136] Optionally, during the adjustment of intermediate representation data, a preview image rendered based on the adjusted intermediate representation data can be displayed in real time on the preview interface of the terminal device, allowing the user to preview the adjustment effect and decide whether to continue adjusting and how to adjust it until satisfied. Of course, the local model can also determine whether the adjusted intermediate representation data meets the display requirements of the target object.

[0137] It should be noted that in some embodiments, after receiving the intermediate representation data from the server, the terminal device can directly execute the subsequent fine-tuning steps without going through the above-mentioned interactive operations. In this regard, after receiving the visual potential field data sent by the server, the terminal device can place the actual target object into the visual potential field to generate a preview image for the user through rendering. The user then determines whether step S51 needs to be executed, or the local model can autonomously determine whether the target object needs to have its position, size, or orientation in the image space adjusted after the intermediate representation data is injected. This application does not limit the determination method.

[0138] Step S52: Based on at least one visual identity feature information in the object feature data, process the corresponding subfield data in the adjusted intermediate representation data to obtain enhanced intermediate representation data that matches the target object.

[0139] Following the above analysis, after obtaining the adjusted intermediate representation data, visual identity feature information from the locally retained object feature data can be combined to process at least one of the corresponding subfield data in the intermediate representation data, such as illumination potential field data, density potential field data, and background structure potential field data. This processing involves operations such as modulation, adsorption, collapse, redirection, and repulsion to obtain enhanced intermediate representation data that matches the target object. The visual identity feature information can include at least one of boundary features, contact area features, material property features, visual centroid features, and protected area features. Different types of visual identity feature information can correspond to different subfield data activation methods, without limitation.

[0140] Optionally, for the boundary features of the target object, a boundary adsorption activation method can be used to process the illumination potential field data and density potential field data; for the contact area features of the target object, a grounding region collapse activation method can be used to process the illumination potential field data (such as shadow and / or reflection components); for the material property features of the target object, a material response activation method can be used to process the illumination potential field data and density potential field data; for the visual center of gravity adjustment of the target object, an attention flow redirection modulation method can be used to process the background structure potential field data and illumination potential field data; for the protected area features of the target object, a protected area repulsion activation method can be used to process the illumination potential field data, density potential field data, and background structure potential field data.

[0141] Based on this, this application can generate edge lighting according to the actual product outline, that is, to form an edge lighting effect that precisely fits the outline near the precise outline of the target object. It can also generate desktop shadows based on the grounding area, ensuring that the final shadow starts from the correct contact point (i.e., the actual boundary between the bottom of the target object and the supporting surface) and naturally diffuses outwards to form a soft shadow effect, rather than starting from a proxy object (such as...). Figure 4The white sphere (a virtual object with reference feature data used to replace the target object) begins to spread from its bottom or an incorrect location. Depending on the material type, metal reflection can be enhanced or fabric reflection reduced. Particle coverage can be suppressed based on the logo area (i.e., a protected area, which may also include other protected areas) to prevent the logo area from being covered or contaminated by background textures, particles, lighting reflections, or structural lines. Background lines can be redirected based on the main visual focus to further enhance the guidance of the viewer's eye, ultimately resulting in an enhanced product display image. This application does not limit the implementation methods of various activation methods.

[0142] In practical applications of this application, multiple activation methods can be pre-configured based on historical data, accumulated expert experience, or other prior knowledge in the field. This can be stored as an activation method library and dynamically updated according to actual circumstances, such as adding new activation methods or adjusting existing activation methods based on changes in domain knowledge. Thus, during the execution of step S52, the local intelligent agent, model, or other processing tools can automatically select an appropriate activation method based on the visual identity feature type contained in the object feature data to process the subfield data in the intermediate representation data.

[0143] Optionally, during the process of determining object feature data, the user can select a suitable activation method and associate it with the corresponding visual identity feature information for storage. Then, when executing step S52, the selected activation method can be directly invoked to process the visual identity feature information associated with the object feature data. Alternatively, during the execution of step S52, the user can select a specific activation method for different types of visual identity feature information. This application does not limit the implementation method of step S52.

[0144] Step S53: Render the target image of the target object based on the enhanced intermediate representation data.

[0145] Since the subfield data, after user interaction adjustment and precise modulation of visual identity feature information, describes how the lighting should be represented, how decorative elements should be distributed, and how the background structure should be organized in the final image, it is input into the local rendering engine or local lightweight generation model for image rendering. For example, the subfield data is decoded into pixel-level RGB images. As needed, the corresponding visual identity feature information can be mapped to the area where the target object is located, and finally a target image that contains the complete identity features of the target object and has the user-customized scene atmosphere can be generated. It can be presented on the interactive interface of the terminal device for users to preview, save or share. If the user is not satisfied with the target image, they can return to step S51 to continue interactive adjustment. After iterative processing, a satisfactory target image is obtained. The implementation process is not described in detail in this application.

[0146] In some embodiments, during the rendering and generation of the target image, the original image data of the target object can be combined with the descriptive data contained in the enhanced intermediate representation data to generate a pixel-level target image. Optionally, this application can be implemented using a layered overlay method, that is, each visual enhancement layer corresponding to the enhanced intermediate representation data is sequentially overlaid onto the original image of the target object to form the final enhanced image. In this regard, after the terminal device obtains the original image data of the target object, such as a product photo or design drawing uploaded by a user, it can generate corresponding visual enhancement layers based on the subfield data in the enhanced intermediate representation data.

[0147] Reference Figure 6 The diagram shown illustrates enhanced intermediate representation data, which may include a background layer, rim light layer, shadow layer, reflection layer, material enhancement layer, decorative detail layer, and atmosphere layer that are precisely matched to the target object. Figure 6 Not all layers are shown. Understandably, illumination enhancement layers, such as highlight layers, shadow layers, reflection layers, and ambient light layers, can be generated based on illumination potential field data; decorative enhancement layers, such as particle layers, texture layers, and line layers, can be generated based on density potential field data; and background enhancement layers, such as platform layers, spatial hierarchy layers, and structural line layers, can be generated based on background structural potential field data. Then, these visual enhancement layers can be sequentially overlaid onto the original image according to a preset stacking order, such as background layer first, then decorative layer, and finally illumination layer. Image compositing methods, such as alpha blending, additive blending, and multiplicative blending, are then used to generate the final enhanced image, i.e., the target image.

[0148] In some embodiments, this application may also employ a rendering pipeline approach, inputting the subfield data from the enhanced intermediate representation data as rendering parameters to the local rendering engine. The rendering engine then generates a target image with realistic lighting, materials, and spatial sense based on the Physically Based Rendering (PBR) pipeline. Optionally, this application may also employ a deep learning rendering approach, inputting the enhanced intermediate representation data and the original image data of the target object into a locally deployed neural network rendering model (i.e., a lightweight generative model), which then generates the target image end-to-end. It should be noted that this model can learn complex visual mapping relationships to generate images with high realism and stylistic consistency. This application does not limit the model type or its acquisition method.

[0149] In summary, users can intuitively adjust the position, size, and orientation of the target object in the image space through interactive operations. The sub-field data in the intermediate representation data adjusts synchronously with the user's actions, achieving a WYSIWYG interactive experience. Alternatively, users can finely adjust the intermediate representation data at the semantic level using natural language prompts, meeting the personalized creative needs of different users and lowering the barrier to entry for professional rendering software. By applying multiple visual identity features to the corresponding sub-field data, precise matching between the target object's realistic appearance and the server-side scene is achieved, avoiding visual distortion caused by simple overlay. This ensures that the final generated target image highly matches the target object in terms of lighting, materials, shadows, background structure, and core asset protection. Furthermore, since the visual identity features are always processed locally, the server only participates in the generation and adjustment of the intermediate representation data, without accessing any sensitive information, thus improving the information security of image processing. In addition, this application supports multiple adjustments and instant previews, allowing users to explore multiple design schemes in a short time, improving creative efficiency and reliably meeting the high-quality image needs of different scenarios such as product display, character display, and brand visuals.

[0150] The processing methods for each of the five visual identity features listed above will be described below. It should be noted that these five processing methods can be used individually or in any combination, depending on the specific needs of the scenario; this application does not impose any restrictions on this.

[0151] When the visual identity features of an object include boundary features, such as precise contour information like a sequence of contour points or a contour mask, a boundary snapping activation method can be used to snap the corresponding subfield data to the vicinity of the real target object's edge. This generates edge lights, contour glows, edge shadows, or boundary decoration layers that match the real contour of the target object. Figure 7 The diagram shown illustrates boundary adsorption activation, generating realistic edge lighting for the laptop interface. This application can adjust the spatial distribution of illumination potential field data or density potential field data in the intermediate representation data within the corresponding contour region based on the boundary features of the target object. This contour region can be determined based on the boundary features of the target object.

[0152] Optionally, this application can first obtain the precise contour information of the target object from the object feature data, such as contour point sequence or contour mask. In this embodiment, the binary subject mask M(P) of the target object is used as an example for illustration. p represents any pixel coordinate in the image space. M(P)=1 indicates that the pixel belongs to the subject region of the target object, and M(P)=0 indicates that the pixel belongs to the background region. Then, the morphological boundary extraction algorithm (such as Canny edge detection or morphological gradient) calculates the boundary set of the subject mask to obtain the boundary mask of the target object, that is, ∂M= Boundary(M), where ∂M belongs to the set of pixel points (contour region) of the contour edge of the target object. The contour region in the boundary mask is marked as 1, and the other regions are marked as 0.

[0153] Subsequently, a boundary distance field can be constructed based on the contour information of the target object. That is, for each pixel location in the image space, the distance from that pixel to the nearest contour point is calculated. Regions with smaller values ​​in this distance field represent positions closer to the target object's contour, while regions with larger values ​​represent positions farther from the contour. Optionally, this application can use Euclidean distance calculation, where the shortest Euclidean distance from any pixel p to the target object's boundary ∂M can be expressed as:

[0154] (1);

[0155] As can be seen, this distance field reflects the proximity of each pixel to the outline of the target object; that is, the pixel distance at the boundary is 0, and the distance increases with distance from the boundary. Therefore, based on the boundary distance field of the target object, the corresponding sub-field data in the adjusted intermediate representation data can be processed. For example, in areas where the distance field value is less than a first preset threshold (i.e., areas adjacent to the outline), the edge light intensity value at the corresponding position in the illumination potential field data is enhanced, with the enhancement magnitude increasing the closer to the outline. This creates an edge light effect that precisely fits the outline near the target object's precise outline. In areas where the distance field value is less than a second preset threshold (i.e., areas near the outline), the particle density value or texture density value at the corresponding position in the density potential field data is reduced, preventing decorative elements from clinging to the subject's edge and covering it, thus creating natural white space or attenuation areas. The second preset threshold can be the same as or different from the first preset threshold, and can be set according to the actual application scenario.

[0156] Optionally, the adjustment of the edge light intensity enhancement can be based on a Gaussian function exp(-d² / σ²) varying with distance d, where σ represents a preset edge light diffusion coefficient. The density attenuation can be based on a sigmoid function (an activation function, but not limited to this) varying with distance, with the density value approaching 0 at the contour and gradually recovering to its original value after a certain distance from the contour. In another possible implementation, this application also constructs boundary adsorption weights based on the boundary distance field to process illumination and density potential field data, such as constructing a boundary adsorption weight function that decays exponentially with distance:

[0157] (2);

[0158] Where, σ b This is a boundary snapping range control parameter used to adjust the width of the edge light's influence. The closer pixel p is to the target object boundary, the more... When the value approaches 0 and belongs to the background region (i.e., M(p)=0), the boundary adsorption weight... The larger the value of M(p), the closer it approaches 1; conversely, when pixel p is far from the boundary of the target object or belongs to the interior of the target object, i.e., when M(p)=1, the boundary snapping weight is higher. The smaller it is, the closer it is to 0. Therefore, The edge light has a significant value only in the vicinity of the target object, thus confining the edge light to a halo around the target object's outline. Based on this, the edge light layer generated by the terminal device can be represented as:

[0159] (3);

[0160] Where E(p) represents the edge light layer, L(p) represents the illumination potential field data, and L dir (p) represents the illumination direction of the pixel in the illumination potential field, i.e., the light direction vector; n(p) represents the normal direction of the pixel outside the boundary of the target object, L color (p) represents the illumination color or color temperature (color value) of a pixel in the illumination potential field, ensuring that the color of the edge light is coordinated with the scene lighting; The dot product of the light direction and the normal can be represented by the cosine of the angle between the incident light and the surface normal. Only positive values ​​are used to ensure that edge light comes from the front direction. This represents the edge light intensity coefficient, used to control the overall enhancement magnitude. This is used to ensure that edge lighting is enhanced only when the lighting direction aligns with the boundary normal.

[0161] By adjusting the illumination potential field data using the above processing methods, or by superimposing the aforementioned edge light layer onto the illumination potential field data, the illumination trends in the processed illumination potential field data that are unrelated to the target object will not directly affect the entire image. Instead, they will be activated near the boundary of the local real target object, thereby forming an edge light or contour atmosphere layer that is consistent with the contour of the target object. That is, a luminous edge that closely fits the contour shape and whose intensity decreases with distance is formed on the outside of the target object contour, which significantly enhances the separation and three-dimensionality of the target object from the background.

[0162] Optionally, this application may also utilize the aforementioned boundary adsorption weights to attenuate the density potential field data, thereby reducing the density of particles, textures, or lines in the background area near the outline of the target object, preventing decorative elements from intruding into the edge of the main body, and maintaining a clear outline.

[0163] Therefore, this application achieves precise spatial control of edge light through a mathematically derived distance field and exponential decay weights. This means the edge light appears only in the immediate vicinity of the subject, its intensity smoothly decays with distance, and there are no hard edges or abrupt changes. Simultaneously, the color and direction of the light effect remain consistent with the scene lighting, resulting in a natural and realistic visual effect. Furthermore, the same boundary adsorption weights can be reused for the decay of the density potential field, achieving a dual effect of edge enhancement and background purification, allowing the enhancement layer to naturally conform to the target object.

[0164] When the visual identity feature information includes contact area information between the target object and a supporting surface such as the ground, table, display stand, or background plane, a grounding region collapse activation method can be used to constrain the shadow potential field or reflection potential field data in the intermediate representation data to the vicinity of the contact area between the target object and the supporting surface, ensuring that the final generated soft shadows, reflections, and contact dark areas match the contact position of the target. Based on this, this application can adjust the intensity of the component values ​​corresponding to the contact area in the illumination potential field data of the intermediate representation data based on the contact area characteristics of the target object. This application does not limit the implementation method of this adjustment. The contact area can be determined based on the contact area characteristics of the target object.

[0165] Optionally, this application can first obtain the contact area information between the target object and the supporting surface from the object feature data. This can be represented as a contact line or a contact surface area, such as the area where the bottom of the target object intersects with the supporting surface, which is denoted as the grounding region G. This application does not limit the method for determining it. For example, based on the above M(p), the depth information obtained from the depth map or 3D model projection in the object feature data (which can be denoted as Z(p), the bottom contour, or manual annotation, the grounding region G is determined, such as G=ContactRegion(M,Z). It is usually a strip-shaped area where the bottom of the target object intersects with the supporting surface, which is represented in the image space as a set of pixels near the lower edge of the target object mask. This area is the physical source of shadows and reflections.

[0166] Then, based on the contact area information, the components in the illumination potential field data corresponding to shadows or reflections can be modulated. For example, within the contact area, the numerical intensity of the shadow component is set to its maximum value to form the darkest part of the shadow; as the pixel position moves upward (i.e., away from the contact area), the numerical intensity of the shadow component gradually decreases according to a preset decay function. The decay function can be in the form of exponential decay, i.e., I_shadow(y) = I_max × exp(-α × y), where y is the vertical distance from the pixel position to the contact area, and α is a preset decay coefficient. This ensures that the final generated shadow naturally diffuses outward from the correct contact position (i.e., the actual boundary between the bottom of the target object and the supporting surface), forming a soft shadow effect.

[0167] Alternatively, this application can also generate a contact shadow layer by constructing a grounding region collapse weight. The collapse weight function, which decays exponentially with distance, can be:

[0168] (4);

[0169] (5);

[0170] Where, σ g This indicates the extent to which a shadow or reflection spreads outward from the grounded area, and is used to adjust the width of the shadow or reflection spreading outward from the grounded area. This represents the shortest Euclidean distance from pixel p to the ground region G. It reflects the proximity of each pixel to the ground region; pixels within the ground region have a distance of 0, while pixels farther from the ground region have a larger distance. Based on this, when pixel p is located within G, , When pixel p is far from G, Approaching 0. This application controls Strictly constrain the response of shadows or reflections to the vicinity of the grounding area to prevent shadows or reflections from appearing in unreasonable locations.

[0171] Subsequently, if the intermediate representation data returned by the server contains the shadow response potential field S(p), it can be combined with the collapse weights. The contact shadow layer Sh(p) is generated, which can be represented as:

[0172] (6);

[0173] in, This represents the shadow intensity coefficient, used to control the overall depth of shadows; This is a reverse mask used to ensure that shadows are rendered only in the background area, preventing shadows from covering the inside of the target object. A contact shadow layer can be overlaid onto the corresponding area of ​​the background image or lighting potential field to achieve a shadow effect with realistic contact.

[0174] Optionally, for projected shadows that need to reflect directionality, such as long shadows produced by oblique light, the lighting direction L can also be introduced. dir The grounding area is offset in the opposite direction of the illumination to simulate the stretching effect of a projection:

[0175] (7);

[0176] in, L is the projection offset coefficient, used to control the length of the projection; proj Let be the projection vector of the lighting direction onto the support surface. Then, the collapse weights are applied to the offset pixel positions, where the offset shadow layer can be represented as:

[0177] (8);

[0178] In this way, when the server only provides the shadow distribution trend, the local area collapses the shadow potential field to the vicinity of the real ground area according to the contact position of the real target object, so that the shadow is consistent with the landing point of the target object and the relationship between the desktop / ground. The shadow is no longer limited to directly below the ground area, but extends along the opposite direction of the light, forming a projection effect that matches the real lighting conditions.

[0179] Optionally, if the intermediate representation data returned by the server contains a reflection response potential field, a contact reflection layer can be generated by referring to the above-mentioned method for generating a contact shadow layer. Based on this, a reflection or reflection effect matching the bottom contour of the main body can be generated on the support surface near the grounding area, further enhancing the sense of integration between the target object and the scene.

[0180] Therefore, this embodiment, through grounding region collapse weighting, precisely constrains the server-generated generic shadow / reflection response to the vicinity of the contact area between the target object and the supporting surface, avoiding shadows or reflections appearing in unreasonable positions, such as floating in the air or deviating too far from the target object. Simultaneously, by introducing illumination direction offset, it simulates the physical laws of projection changing with illumination direction in the real world, giving the generated shadows directionality and length variations, resulting in a more realistic and natural visual effect. Furthermore, the generation of the contact reflection layer further enhances the immersion and realism of the scene.

[0181] When the visual identity information of an object's feature data includes material attribute features, the optical parameters corresponding to the illumination potential field data and / or density potential field data in the intermediate representation data can be adjusted based on the target object's material attribute features. These parameters, such as the intensity value, frequency distribution, or response amplitude on the corresponding channel, can be adjusted to produce differentiated rendering results for the same server-side visual potential field due to differences in the object's actual material. Material categories can include, but are not limited to, metal, glass, plastic, fabric, leather, ceramic, and wood. Different material categories correspond to different modulation parameters (control commands), using the specular highlight enhancement coefficient k as an example. spec Roughness adjustment coefficient k rough Reflection intensity coefficient k ref and texture enhancement coefficient k tex Using these four types of modulation parameters as examples, this application can change the activation state of the corresponding subfield data in the local visual potential field data by adjusting one or more modulation parameters, so as to visually be equivalent to changing the optical parameters of the material.

[0182] It is evident that the four modulation parameters mentioned above have a close mapping relationship with the rendering physical quantities of reflectivity, diffuse reflectivity, specular concentration, and texture frequency. The terminal device activates the visual potential field data based on the modulation parameters, thereby influencing these rendering physical quantities to achieve material adaptation. Among these, k... spec There is a direct mapping relationship between k and the concentration of highlights; the two are positively correlated. spec The larger the value, the sharper and more concentrated the highlights; the smaller the value, the more diffuse and soft the highlights. ref There is a direct mapping relationship between k and reflectivity; the two are positively correlated. ref The larger the value, the stronger the Fresnel reflection, and the more obvious the mapping of the environment on the surface; k rough There is an indirect mapping relationship with specular concentration and total reflectance, such as k rough Negatively correlated with specular concentration, i.e., k rough The larger the value, the more diffuse the highlights (becoming softer and wider), which also affects the full reflection performance, such as making diffuse reflection more uniform on rough surfaces. tex There is a direct mapping relationship between k and texture frequency; the two are positively correlated. texThe larger the value, the higher the density and frequency of the surface texture. Based on this, the illumination, reflection, specular, or texture potential fields in the visual potential field data returned by the server or after preliminary adjustment can be locally adjusted. For example, reflection and specular highlights can be enhanced for metallic materials, specular reflection can be reduced for plastic materials, and transparent reflection or edge refraction effects can be increased for glass materials.

[0183] In some embodiments, this application can flexibly adjust / configure the values ​​of the above modulation parameters according to material characteristics and visual effect requirements. Taking metal, plastic, glass, and fabric as examples, the metal material category can be configured with a higher k value. spec and k ref And configure a lower value for k. rough and k tex , such as k spec =1.5, k ref =1.8, k rough =0.2, k tex =0.1; The plastic material category can be configured with a medium value for k. spec k rough and k tex And configure a lower value for k. ref, such as k spec =0.8, k rough =0.5, k tex =0.5, k ref =0.3; The glass material category can be configured with a higher value for k. ref The value of k is moderate. spec and the lower value of k rough and k tex , such as k spec =0.9, k rough =0.3, k tex =0.2, k ref =1.5; the fabric material category can be configured with a higher value for k. rough and k tex and the lower value of k spec and k ref , such as k spec =0.2, k rough =0.9, k tex =1.6, k ref =0.1, but is not limited to the value example in this application, and can be flexibly adjusted according to the actual situation.

[0184] Based on this, this application can obtain the material category map of the target object from the object feature data, where T(p) represents the material category label at pixel position p. Optionally, this application can analyze the target object image to determine its material category using a locally pre-trained material classification model, a general image processing model, or a lightweight AI model; alternatively, the user can manually label the material category corresponding to different pixel positions. This application does not limit the method for identifying the material category at each pixel position. Then, the modulation parameter corresponding to T(p) of the target object can be found from the pre-configured mapping relationship of modulation parameters corresponding to different material categories (such as a mapping table, the storage method of which is not limited in this application). spec and k ref The reflection / spectral response potential field data R(p) and illumination potential field data L(p) in the intermediate representation data are weighted and fused to obtain the reflection enhancement layer Re(p):

[0185] (9);

[0186] Where M(p) is the target object mask, ensuring that the reflection enhancement only applies to the target object region; This setting controls the intensity of the specular highlight, such as a stronger highlight for metallic materials and a weaker highlight for fabric materials; This setting controls the intensity of reflection; for example, metal and glass materials receive stronger environmental reflection, while plastic and fabric materials receive weaker reflection.

[0187] Similarly, based on the modulation parameter k corresponding to T(p) tex The density potential field data D(p) in the intermediate representation data is processed to obtain the texture enhancement layer Te(p):

[0188] (10);

[0189] Where D(p) represents the channel data related to texture density in the density potential field data. Used to control the degree of surface texture display, such as giving fabric materials a stronger texture enhancement (e.g., fabric weave patterns), and metal and glass materials a weaker texture enhancement, making their surfaces smooth, etc.

[0190] Next, the reflection enhancement layer and texture enhancement layer can be combined according to the fusion coefficient to obtain the final material enhancement result Mat(p). This result is then superimposed onto the corresponding subfield data in the intermediate representation data. For example, the reflection enhancement layer can be superimposed onto the specular reflection channel or independent high-potential field of the illumination potential field data, and the texture enhancement layer can be superimposed onto the texture density channel of the density potential field data. This ensures that the illumination reflection and surface texture of the target object region in the obtained enhanced intermediate representation data have been differentially modulated according to the real material. Mat(p) can be represented as:

[0191] (11);

[0192] in, and These are the global fusion coefficients for the reflection enhancement layer and the texture enhancement layer, respectively. They are used to adjust the relative weight of the two types of enhancement effects and can be determined or dynamically adjusted according to the target object type and visual effect requirements. , wait.

[0193] As can be seen, this application achieves adaptive material adjustment of the server's general visual potential field by driving modulation parameters through a local material category map. Using this method, even if the server feeds back the same visual potential field data, the terminal device will produce different activation effects for objects of different material categories locally. That is, the activation results of different material category regions of the target object are different; for example, metal areas produce more obvious highlights and reflections, plastic areas produce softer diffuse reflections, and glass areas produce stronger transmission or edge brightness changes. In this way, material-matched rendering effects can be obtained without the server needing to know the actual material properties of the target object, thus protecting the potential trade secret of the target object's actual material properties while fully utilizing the server's powerful scene generation capabilities.

[0194] For example, suppose the target object is a smartphone that includes both a metal frame and a plastic back panel, and the k of its metal frame area... spec A higher value produces bright and sharp highlight bands; k ref The higher the value, the more the reflection of the surrounding environment is projected onto the border; k tex The value is low, and the surface remains smooth and textureless. k in the plastic backplate area. spec Using a low to medium value produces soft diffuse highlights; k ref Take a low value; there is no obvious environmental reflection. tex By taking the median value, a subtle frosted texture is revealed. As a result, in the final rendering, the metal frame shines with sharp highlights and ambient reflections, while the plastic back panel presents soft diffuse reflections and a subtle frosted texture. The two exhibit distinctly different material properties under the same lighting conditions, resulting in a realistic and natural visual effect.

[0195] Optionally, this application can also flexibly configure the range of optical parameters such as reflectance, diffuse reflectance, specular concentration, and texture frequency for various materials according to material characteristics and visual effect requirements, without limitation. Taking metal, plastic, glass, and fabric as examples, the reflectance of metal can be configured with a higher level, such as a value in the range of 0.7-0.9; the diffuse reflectance can be configured with a lower level, such as a value in the range of 0.1-0.2; the specular concentration can be configured with a roughness in the range of 0.1-0.3; and the texture frequency can be configured with a higher level. The reflectance of glass can be configured with a medium-to-high level, such as a value in the range of 0.5-0.8; the diffuse reflectance can be configured with a lower level, such as a value in the range of 0.05-0.15; the specular concentration can be configured with a roughness in the range of 0.05-0.2; and the texture frequency can be configured with a lower level. For the plastic material category, the reflectivity can be configured to a medium level, such as 0.3-0.5; the diffuse reflectivity can be configured to a medium level, such as 0.4-0.6; the specular concentration can be configured to a medium level, such as roughness in the range of 0.3-0.5; and the texture frequency can be configured to a medium level. For the fabric material category, the reflectivity can be configured to a low level, such as 0.1-0.2; the diffuse reflectivity can be configured to a high level, such as 0.7-0.9; the specular dispersion can be configured to a high level, such as roughness in the range of 0.7-0.9; and the texture frequency can be configured to a high level.

[0196] Based on this, the terminal device can determine the corresponding optical parameter set according to the material category of the identified target object, and then use this optical parameter set to process the illumination potential field data in the adjusted intermediate representation data. This involves adjusting the illumination intensity value at the corresponding position in the illumination potential field data; for example, increasing the reflected light intensity of metallic materials and decreasing the reflected light intensity of fabric materials. It can also adjust the reflection intensity value; for example, increasing the specular reflection component of metallic materials and decreasing the specular reflection component of fabric materials. Depending on actual needs, the density potential field data in the adjusted intermediate representation data can also be processed according to the texture frequency parameter corresponding to the material category. This involves adjusting the texture frequency value at the corresponding position in the density potential field data; for example, increasing the texture frequency of metallic materials to present a brushed or frosted effect, increasing the texture frequency of fabric materials to present a woven texture effect, and decreasing the texture frequency of glass materials to present a smooth surface effect.

[0197] When the visual identity feature information of an object includes visual centroid features, the spatial guidance distribution corresponding to the background structural potential field data or the illumination gradient direction corresponding to the illumination potential field data in the intermediate representation data can be adjusted based on the visual centroid features of the target object. Here, the visual centroid refers to the position of the target object that is most visually attractive. For example, for a smartphone, the visual centroid could be the main camera area; for a car, it could be the front grille area; and for a character IP, it could be the character's facial area.

[0198] In this way, after obtaining the visual center of gravity of the target object from the object feature data, the background structural potential field data in the intermediate representation data, which has been initially adjusted by the server feedback or through the above interactive operations, can be processed accordingly. For example, the structural direction vectors of each pixel position in the background structural potential field data can be adjusted so that background structural lines, such as decorative lines and texture extension directions, converge towards the visual center of gravity. For example, for each pixel position p (which can be represented by pixel coordinates (x, y)), the direction vector v from pixel position p to the visual center of gravity position c is calculated, i.e., v = c - (x, y). Then, this direction vector is weighted and fused with the original structural direction vector of position p in the background structural potential field data, so that the background structural direction gradually shifts towards the visual center of gravity.

[0199] Optionally, this application can also process the illumination potential field data in the adjusted intermediate representation data based on the visual center of gravity position, such as adjusting the direction of the illumination gradient in the illumination potential field data so that the transition direction of light and shadow also points to the visual center of gravity position, further enhancing the guiding effect on the viewer's gaze. After the above processing, the background lines and the light and shadow gradient together guide the viewer's gaze naturally to the core area of ​​the target object, improving the focusing effect of the display.

[0200] In one possible implementation, the visual centroid position c of the target object can be calculated using the following formula:

[0201] (12);

[0202] Where p represents the pixel coordinates in the image space, M(p) is the binary subject mask of the target object, M(p)=1 indicates that the pixel position belongs to the subject region of the target object, i.e., the target object region, and M(p)=0 indicates that the pixel position belongs to the background region of the target object, in which case the pixel is a background pixel. This application can obtain the geometric centroid position of the target object in the image space, i.e., the visual centroid position c, by taking the weighted average of the pixel coordinates of all pixels within the subject region.

[0203] Next, a unit direction vector v(p) can be constructed pointing from any background pixel to the visual centroid position c, and attention retargeting weights can be constructed based on the distance from the pixel position to the target object region:

[0204] (13);

[0205] (14);

[0206] Here, normalize represents the normalization operation. Let p be the shortest Euclidean distance from pixel p to the target object mask M. This parameter controls the scope of the redirection effect. Its weight is larger (close to 1) when it is near the target object and approaches 0 when it is far away from the target object, ensuring that the redirection effect is mainly concentrated in the background area around the target object.

[0207] Based on this, this application can redirect the attention flow field A(p) in the intermediate representation data to obtain a new attention flow field. In this context, the attention flow field can be used as a derived component of the background structural potential field to describe the directional distribution of visual guidance.

[0208] (15);

[0209] in, This represents the attention redirection strength coefficient, which controls the strength of the redirection; The weights calculated above are used to control the spatial influence range. The direction vector points towards the visual center of gravity. Equation (15) indicates that in the region near the target object, the original attention flow field is gradually replaced with the direction pointing towards the visual center of gravity; in the region far from the target object, the original attention flow field remains unchanged. The normalization operation ensures that the output is a unit direction vector.

[0210] Similarly, this application can redirect the background structural potential field data B(p) (which contains spatially guided distribution information) in the intermediate representation data to obtain redirected background structural potential field data. :

[0211] (16);

[0212] in, This represents the background structure redirection intensity coefficient. After redirection, the perspective lines, decorative trends, and spatial gradients of the booth in the background all converge towards the visual center of gravity of the target object. It is evident that this embodiment, by redirecting the server-generated general attention flow and background structure towards the locally calculated visual center of gravity of the target object, allows the visual guide lines, light and shadow directions, and decorative trends of the final image to naturally converge towards the target object, significantly enhancing the prominence and structure of the target object. Figure 1 Consistency. In this way, even if the server does not know the exact location of the target object when generating intermediate representation data, and only knows a rough placeholder (target space attribute information), the local redirection operation can make the background layout accurately adapt to the visual center of gravity of the real target object.

[0213] When the visual identity feature information of the object feature data includes protected area features, the response intensity of the illumination potential field data, density potential field data, and / or background structure potential field data in the intermediate representation data within the protected area is adjusted based on the protected area features of the target object. This prevents the protected area in the final rendered target image from being occluded / covered by particles, textures, decorative lines, light spots, or background elements. The protected area refers to the area on the target object that requires special protection and should not be covered or interfered with by any visual enhancement elements. Examples include brand logo areas, brand identification / text areas, key structural areas, and core functional areas such as interfaces, buttons, or heat dissipation holes. These areas can be automatically identified by the local model or manually labeled by the user. This application does not limit the type or identification method of the protected area.

[0214] Based on this, this embodiment can obtain protected area features such as the protected area location information of the target object from the object feature data. The protected area location information can be represented as a protected area mask (pixel value of 1 within the protected area and 0 outside the area) or the boundary outline of the protected area. Then, the subfield data in the intermediate representation data returned by the server or after preliminary adjustment can be processed accordingly. For example, within the protected area, the illumination intensity value at the corresponding position in the illumination potential field data is multiplied by a weighting coefficient less than 1 (such as 0.2) or directly set to zero to avoid overexposure, reflection, or highlight interference in the protected area; at the same time, the chromaticity value in the illumination potential field data is adjusted to a neutral color or a preset protective color to avoid color pollution.

[0215] Optionally, within the protected area, the density values ​​at corresponding locations in the density potential field data can be set to 0 or extremely low values ​​to ensure that decorative elements such as particles, textures, and lines do not cover the protected area. Furthermore, a gradient approach (e.g., gradually transitioning the density value from 0 to the original value) can be used at the edge of the protected area to ensure a smooth and natural boundary transition. Additionally, the spatial hierarchy or structural orientation values ​​at corresponding locations in the background structural potential field data can be adjusted to prevent background structural lines from crossing or interfering with the protected area, ensuring the clarity and information integrity of the protected area in the final generated target image.

[0216] In one possible implementation, this application generates a protected region mask Q(p) for the target object using a target detection model or manual annotation. Q(p) = 1 indicates that pixel p belongs to the protected region, and Q(p) = 1 indicates that pixel p does not belong to the protected region. To ensure a smooth transition of the suppression effect at the edge of the protected region and avoid hard edges or abrupt changes, the protected region mask is Gaussian blurred to obtain the expanded protected region mask. Its generation method can be expressed as:

[0217] (17);

[0218] in, The standard deviation of the Gaussian kernel is used to control the range of expansion. The expanded range... The value approaches 1 within the protected area and smoothly decays to 0 outside the protected area, forming a gentle transition zone. Based on... Protection gating functions can be constructed:

[0219] (18);

[0220] in, To protect the strength coefficient, when When external decorative elements are completely suppressed within the protected area, the gating value is 0; when At this time, a small amount of low-intensity visual effect is allowed to be retained, in which case the gating value is greater than 0. The gating function takes a lower value inside the protected area and a value close to 1 outside the protected area.

[0221] For the density potential field data D(p), background structure lines B(p), and attention flow field data A(p) within the protected area of ​​the visual potential field data, the same gating value can be used for suppression, i.e.:

[0222] (19);

[0223] (20);

[0224] (twenty one);

[0225] In this way, the operation represented by the above formula (19) significantly reduces or even reduces the particle density, texture density, and line density to zero within the protected area, ensuring that the protected area, such as the logo and text, is not covered by background decorative elements. The booth lines, perspective structure, and visual guidance direction in the background are also suppressed within the protected area to prevent structural elements from passing through or covering sensitive areas.

[0226] In summary, this embodiment achieves precise local suppression of decorative elements, background structures, and visual flow returned by the server through a protective region mask and gating functions. The protected region is dynamically generated locally based on the target object image, without server intervention, thus protecting sensitive visual assets while maintaining the integrity of background decorations in other areas. Gaussian spreading and smooth transitions avoid hard edges or visual breaks at the suppression boundaries, allowing the protected region to blend naturally with the surrounding background.

[0227] For example, if the target object is a smartphone with a brand logo and model number printed on the back, a camera module at the top, and volume buttons and a SIM card slot opening on the side, a protective area mask Q(p) is generated locally using OCR (Optical Character Recognition) and keypoint detection to cover all the above-mentioned protective areas. The density potential field data returned by the server generates dense, tech-inspired light trails and flickering light spots around the smartphone. After protective area suppression, the density potential field values ​​of the logo area and the model number area are set to zero, and the light trails and light spots completely disappear, making the logo and model number clearly visible. The density potential field value of the camera module area is significantly reduced, retaining only a very low-density faint halo, without affecting the appearance of the camera. The density potential field value of the volume button area is suppressed, and the button outline is not obscured by decorative elements, thus ensuring that all core visual assets remain intact and recognizable in the final rendered target image, while background decorations are rendered normally in other areas.

[0228] It should be noted that, regarding the activation method in which the terminal device uses locally stored object feature data to adaptively activate the intermediate representation data returned by the server or implemented through interactive operations after preliminary adjustment, so that the obtained intermediate representation data matches the visual characteristics of the local real target object, thereby achieving both privacy protection and rendering instructions, in addition to the activation methods listed above, other activation methods can also be used, as long as they can achieve the activation purpose, they are all within the scope of protection of this application.

[0229] Optionally, activation methods based on depth / height features can be employed. This involves adjusting the illumination potential field in the intermediate representation data based on the 3D depth information or height map of the target object. For example, the light intensity distribution in the illumination potential field data can be modulated according to the height difference between different areas of the target object's surface, resulting in stronger highlight response in raised areas and natural self-occluding shadows in recessed areas. Alternatively, activation methods based on color / gamut features can be used. This modulates the color channels of the illumination potential field in the intermediate representation data according to the main color distribution of the target object, ensuring that the color temperature and hue of the scene lighting are coordinated with the main color. For example, a warm-colored main subject is paired with warm-colored ambient light, and a cool-colored main subject with cool-colored ambient light, to avoid color clashes. Finally, activation methods based on semantic segmentation features can be used. This involves semantically segmenting the target object to identify different component categories, such as the screen area, bezel area, camera area, and button area of ​​a mobile phone. Different activation parameters can then be applied to different component categories to achieve fine-grained potential field adjustment at the component level.

[0230] Furthermore, activation methods based on user preference features can be employed. This involves learning personalized activation parameter combinations based on at least one of the user's historical preference data, such as preferred lighting style, decoration density, and color tendency, allowing the server-side general visual potential field data to be customized according to user preferences. Activation methods based on environmental context features can also be used. This involves activating the intermediate representation data to match the environment based on the real-world context of the target object, such as indoor / outdoor, day / night, natural / urban, etc. For example, if the original image of the target object is identified as being taken in a beach environment, the blue tone and natural light representation can be enhanced based on the visual potential field data returned from the server. Alternatively, a machine learning-based activation method can be used. This involves deploying a lightweight neural network model locally, inputting object feature data and server-side intermediate representation data into the model, and directly outputting enhanced intermediate representation data through forward inference. The network parameters of this model can be obtained through offline training. The training objective is to make the activated potential field as close as possible to the actual rendering result. This application does not restrict the training implementation method of this model.

[0231] In video scenarios, activation methods based on motion / dynamic features can also be used. This involves smoothing the intermediate representation data temporally based on motion vectors or optical flow information between adjacent frames, ensuring temporal consistency in visual potential field data adjustments across consecutive frames and avoiding flickering or jumps. When multiple target objects exist in the image space, a multi-subject collaborative activation method can be employed. This involves performing multi-subject collaborative activation on the intermediate representation data based on the relative positions, hierarchy, and occlusion relationships among multiple target objects. For example, a primary object can be identified and its visual center of gravity can be used to redirect attention flow, while ensuring that secondary objects also receive appropriate potential field responses.

[0232] In practical applications, the activation methods listed above can also be combined in a series, parallel, or conditionally triggered manner. For example, boundary feature activation can be performed first to enhance edge light, followed by material property activation to modulate optical parameters, and finally protection zone feature activation can be performed to ensure that the core area is not covered. The execution order and combination strategy of each activation method can be flexibly configured according to the specific scenario, and this application does not impose any restrictions on this.

[0233] Reference Figure 8 This is a flowchart illustrating an image processing method proposed in Embodiment 4 of this application. Applied to the server side, it works in conjunction with the aforementioned terminal-side method to form an end-to-cloud collaborative image processing system. Figure 8 As shown, the image processing methods executed by the server may include:

[0234] Step S81: Receive reference feature data of the target object sent by the terminal device.

[0235] The server receives reference feature data sent by the terminal device via a communication network. This reference feature data is generated locally by the terminal device based on the object feature data of the target object. It is used to provide the server with the input information required to generate intermediate representation data. It can be used to characterize the target object's spatial attribute information in the image space, but does not contain visual feature information that can be used to reconstruct the target object's visual identity. Therefore, even if the reference feature data is intercepted by a third party during transmission, it cannot be used to identify or imitate the target object's core visual assets, ensuring the security of the target object's visual identity information. The process of generating the reference feature data can be referred to the description in the corresponding section of the terminal-side embodiment above; it will not be repeated here.

[0236] Step S82: Generate intermediate representation data based on reference feature data. This intermediate representation data is used to characterize the distribution trend of visual elements in the image space.

[0237] Step S83: Send intermediate representation data to the terminal device so that the terminal device generates a target image of the target object based on the object feature data and intermediate representation data of the target object. The object feature data contains visual identity feature information of the target object.

[0238] Based on the above analysis, intermediate representation data is a dataset used to characterize the distribution trends of visual elements in image space. It transmits the server-side model's understanding of the image's visual effects to the terminal device in data form, allowing the terminal device to perform subsequent processing and rendering locally. The server can generate intermediate representation data, such as the aforementioned visual potential field data, based on the input reference feature data using a pre-trained dedicated model (i.e., the server-side model).

[0239] In the training process of the server-side model, corresponding sample reference feature data, sample potential field data, and sample images can be constructed for different sample objects. The initial model to be trained generates corresponding intermediate representation data for each sample reference feature data. Furthermore, the activation method described above can be used to process the intermediate representation data. Based on the processed intermediate representation data and sample object feature data, a sample object image is rendered. Then, based on the potential field loss between the intermediate representation data and the sample potential field data, and the image loss between the sample object image and the sample image, the network parameters of the initial model can be adjusted. The model with the adjusted network parameters is then iteratively adjusted until the training termination condition is met, such as reaching a preset number of iterations or loss convergence. The final model is then used as the server-side model to generate intermediate representation data for the target object. Optionally, this application can jointly train the aforementioned server-side model and the local model on the terminal device. The training process is similar and will not be detailed here.

[0240] In this application, the intermediate representation data returned by the server does not contain visual identity feature information of the target object, but only provides the distribution trends of scene illumination, density, and background structure. The terminal device uses the visual identity feature information held locally to activate and adjust the intermediate representation data, and finally generates a target image that contains both complete identity features of the target object and professional-grade scene effects. This generation process can be referred to the description of the corresponding part of the terminal-side embodiment above, and will not be repeated here.

[0241] In summary, the server only accesses the anonymized reference feature data throughout the entire processing, and cannot determine the visual identity of the target object, thus fundamentally eliminating the risk of privacy leakage. Leveraging its powerful computing capabilities, the server rapidly generates high-quality intermediate representation data based on the target's spatial attribute information, providing professional scene layout solutions for terminal devices. Furthermore, the server completely decouples the generated intermediate representation data from the target object's identity, ensuring both the matching of the scene layout with the target object's occupancy in the image space and preventing the leakage of the target object's visual identity information.

[0242] Reference Figure 9 This is a flowchart illustrating an image processing method proposed in Embodiment 5 of this application. This method is applicable to the server side. This embodiment describes a possible implementation method for generating intermediate representation data based on reference feature data, such as... Figure 9 As shown, the implementation method may include:

[0243] Step S91: Receive reference feature data of the target object and rendering requirement information for the target object sent by the terminal device.

[0244] In this application, the reference feature data of the target object can exist in the form of a proxy such as an abstract geometric body, a fuzzy placeholder, or a non-ontology substitute form, to provide a spatial reference for the subsequent generation of visual potential field data, that is, to determine the approximate occupying area, center position, outer range, orientation trend, and occlusion level of the proxy in the image space. The content and generation method of the reference feature data can be referred to the description of the corresponding part of the above end-side embodiment, and will not be repeated here.

[0245] In this embodiment, the target spatial attribute information of the target object and the rendering requirement information together constitute the reference feature data. The rendering requirement information may include at least one of the following: target styles such as cyberpunk, minimalism, vaporwave, realism, studio, outdoor nature, science fiction, or retro; lighting directions such as upper right 45 degrees, top lighting, side backlighting, softbox, and backlighting; color atmospheres such as warm orange tones, cool blue tones, high saturation, black and white, and Morandi color schemes; background types such as gradient background, blurred background, geometric background, indoor / outdoor landscape, and natural scene; decorative density such as sparse, moderate, and dense, or numerical density parameters; display angles such as 30 degrees downward, eye level, and upward; aspect ratios such as 16:9, 4:3, and 1:1; and output resolutions such as 512×512, 1024×768, 1024×1024, 1920×1080, and 2048×2048. This information is transmitted from the terminal device to the server in a structured data format.

[0246] Optionally, users can input the above rendering requirements through the interactive interface of the terminal device. The input methods may include, but are not limited to: drop-down menu selection, slider adjustment, text input, or natural language description, such as "I want a product display image with warm tones, a blurred background, and technological lighting effects." This application does not restrict the input method for rendering requirements.

[0247] Step S92: Obtain the target spatial attribute information from the reference feature data.

[0248] The server can parse and extract the target spatial attribute information of the target object in the image space from the received reference feature data. That is, the coarse-grained place description obtained by the terminal after desensitizing and dimensionality reduction of the original spatial attribute information is only used to determine the approximate location of the target object in the picture, the area it occupies, and its orientation, without containing any precise geometric details that can restore its identity. This application does not restrict the content and generation method of the target spatial attribute information.

[0249] Step S93: Based on the output resolution in the rendering requirement information, establish a standardized image space grid.

[0250] Step S94: Based on the target spatial attribute information and rendering requirement information, generate illumination potential field data, density potential field data and background structure potential field data on the standardized image spatial grid, respectively, as intermediate representation data adapted to the reference feature data.

[0251] The server can create a standardized H×W image space grid in the logical space based on the output resolution specified in the rendering requirements. Here, H represents the number of grid cells in the height direction of the image space, and W represents the number of grid cells in the width direction. The values ​​of H and W are determined by the output resolution in the rendering requirements. For example, if the output resolution is 1024×1024, then H=1024 and W=1024, which can be in pixels, with each grid cell representing one pixel.

[0252] Each network unit can be assigned two-dimensional coordinates in the image plane, i.e., spatial coordinates within the range of [0, W-1] × [0, H-1]. The offset of the unit relative to the target object's occupied area, i.e., the relative position, is used for subsequent radial distribution calculations. The connectivity relationship between the unit and its neighboring units, such as four-neighbor or eight-neighbor relationships, is used for gradient calculation and spatial smoothing. A normalized index can also be constructed for each unit, that is, the spatial coordinates are mapped to the [0, 1] interval, which is convenient for model processing.

[0253] As can be seen, the standardized image spatial grid serves as a common coordinate system for the generation of all subsequent visual potential fields, ensuring that the three sub-fields of illumination, density, and background structure are generated, aligned, and combined within a unified coordinate system. In other words, the standardized image spatial grid provides a unified spatial plane for the generation of subsequent sub-field data. The illumination potential field data, density potential field data, and background structure potential field data are all generated on this grid, and the data are aligned within the same spatial coordinate system, ensuring that the sub-field data remain spatially consistent during subsequent local processing. It should be noted that the generation processes of the three sub-field data are independent of each other, but they share the same spatial reference system, i.e., the common coordinate system.

[0254] Based on the descriptions of the three sub-field data in the above embodiments, one possible implementation for generating illumination potential field data is to use the center position and outer radius of the target object as a reference, combined with the lighting direction and color atmosphere in the rendering requirements, to generate illumination potential field data. Optionally, the global master ray incident direction is determined according to the lighting direction in the rendering requirements. Then, with the center of the target object as the origin, the offset angle relative to the master ray direction is calculated for each grid cell, forming a light distribution field. A Gaussian decay function is applied with the target object's occupied area as the center to obtain the light intensity distribution, i.e., grid cells closer to the occupied area receive higher illumination intensity, while those farther away gradually decay to the ambient light level. The decay radius is proportional to the outer radius of the target object. Alternatively, according to the color atmosphere in the rendering requirements, a corresponding color temperature and color value can be assigned to each grid cell, and a color gradient can be set in the image, thereby outputting illumination potential field data with an H×W×3 data structure.

[0255] In another possible implementation, this application can obtain the user-specified lighting direction and color atmosphere based on rendering requirement information, and convert the user preference into specific light source parameters. For example, if the user selects "warm light at 45 degrees to the upper right," the server determines that the main light source is located at the upper right of the image space, i.e., the light source direction vector points to the upper left. The color temperature is warm, such as a color temperature value of approximately 3500K, corresponding to an orange-red RGB chromaticity. Then, for each grid cell on the standardized image space grid, the server calculates the direction vector and distance relative to the center position of the target object indicated by the target space attribute information. The server calculates the light intensity received by the grid cell based on the angle between the main light source direction and the direction vector. For example, if the angle between the direction vector and the light source direction vector is less than 90 degrees, it means that the grid cell is located on the side of the target object facing the light source, and the grid cell receives a higher light intensity; if the angle between the direction vector and the light source direction vector is greater than 90 degrees, it means that the grid cell is located on the side of the target object away from the light source, and the grid cell receives a lower light intensity, being in a shadow area. The light intensity value can be calculated according to the cosine law or a more complex lighting model, and this application does not impose any restrictions on it.

[0256] Next, chromaticity values ​​are assigned to each grid cell according to the color atmosphere specified by the user. For example, in a "warm orange" atmosphere, well-lit areas (high intensity) are set with higher R values ​​and lower G and B values ​​to present a warm orange hue, while shadow areas are set with cooler complementary colors (such as bluish-purple), creating a visual effect of warm and cool contrast. In a "cool blue" atmosphere, well-lit areas are set with cool blue values, such as lower R and G values ​​and higher B values, while shadow areas are set with darker cool colors. The server can encode the calculated light direction, intensity, and chromaticity information into a multi-dimensional array and store it in the corresponding grid cell of the standardized image space grid as illumination potential field data.

[0257] During the generation of density potential field data, the server can use the target object's occupied area as a reference, combined with the decoration density level specified in the rendering requirements. To this end, the standardized image space grid can be divided into three concentric regions: the target object's occupied area (which can be called the inner region), a ring-shaped band extending outwards from the target object's occupied area by a certain number of pixels (which can be called the transition region), and the remaining portion of the image (which can be called the outer region). The density value of the inner region is set to 0 to ensure that decorative elements such as particles, textures, or lines do not cover the target object itself; the density value of the transition region increases linearly from the inside out, forming a smooth transition; the density value of the outer region is set to the density level specified in the rendering requirements. Low / medium / high correspond to different baseline values: sparse density baseline values ​​are set to lower values, such as 0.2; moderate density baseline values ​​are set to middle values, such as 0.5; and dense density baseline values ​​are set to higher values, such as 0.8.

[0258] Furthermore, this application can determine the type of decorative elements based on the target style in the rendering requirements information. Different types of decorative elements can be distinguished in the density potential field data through different channels (such as particle density, texture density, and line density) or different numerical distributions. For example, cyberpunk style corresponds to neon lines, mesh textures, and floating data points; minimalist style corresponds to simple geometric shapes and fine lines; vaporwave style corresponds to retro pixels and glitch art textures; futuristic style corresponds to glowing lines and particle streams; and natural scene style corresponds to scattered petals, light spots, or fog particles. The density values ​​of each mesh unit calculated above are encoded into a multidimensional array and stored in the corresponding mesh unit of the standardized image space grid as density potential field data.

[0259] In the process of generating background structural potential field data, the outer radius and orientation trend of the target object can be used as a reference, combined with the background type and display angle in the rendering requirements. Optionally, based on the outer radius of the target object, a platform or support surface can be planned directly below the target object's occupied area. The size of the platform matches the outer radius, and its position is adjusted according to the display angle. The standardized image space grid is divided into spatial layers such as foreground layer, subject layer, background layer, and infinity layer. The subject layer corresponds to the target object's occupied area. The background layer and foreground layer are assigned values ​​according to the depth gradient. For example, the area where the target object is located and its immediate vicinity are divided into the foreground layer, with a spatial layer value close to 0. The area within a certain distance behind the target object is divided into the midground layer, with a spatial layer value of approximately 0.5, which is usually used to place the platform or support structure. The area beyond the midground layer is divided into the background layer, with a spatial layer value close to 1, which is used to place background elements.

[0260] Optionally, this application can divide the spatial hierarchy based on a distance threshold method. That is, for each network unit, the nearest distance from the calculator to the target object's occupied area is used. If the distance is less than a first threshold, the network unit is determined to belong to the foreground layer; if the distance is between the first threshold and the second threshold, the network unit is determined to belong to the midground layer; if the distance is greater than the second threshold, the network unit is determined to belong to the background layer.

[0261] This application can also determine the geometric structure of the background based on the display angle and background type in the rendering requirements. For example, if the display angle is 30 degrees overhead, a flattened elliptical booth area (located in the mid-ground layer) will be generated in the background structural potential field data, with the background behind the booth contracting into the distance, creating a overhead perspective effect. If the display angle is eye-level, a rectangular booth area (located in the mid-ground layer) will be generated in the background structural potential field data, with the background extending horizontally, creating an eye-level composition effect. If the display angle is upward, an upward-extending booth area will be generated in the background structural potential field data, with the background contracting upward. If the background type is a gradient background, no obvious booth structure will be generated in the background structural potential field data; instead, a smooth spatial gradient will be generated. Optionally, this application can also generate a perspective mesh or structural vector field based on the background type. For example, an indoor scene generates perspective lines with vanishing points, while an outdoor scene generates a horizon and a skyline.

[0262] Furthermore, this application can also determine the extension direction of background structural lines based on the orientation trend in the target spatial attribute information. If the target object's orientation trend is to the right, the background structural lines converge to the right, creating a rightward perspective; if the target object's orientation trend is centered and facing forward, the background structural lines radiate outwards from the center, creating a symmetrical spatial sense; if the target object's orientation trend is to the upper left, the background structural lines converge to the upper left. Finally, the calculated spatial hierarchy values ​​and structural directions are encoded into multidimensional arrays and stored in the corresponding grid cells of the standardized image spatial grid as background structural potential field data.

[0263] The server aligns the generated illumination potential field data, density potential field data, and background structure potential field data on a standardized image space grid and encapsulates them into intermediate representation data. These three sub-fields share the same H×W spatial dimension and are each stored independently as a multi-dimensional array, collectively forming a multi-channel visual potential field data packet. This packet serves as intermediate representation data adapted to the reference feature data and is sent to the terminal device via the communication network. Upon receiving this data, the terminal device can read the illumination potential field data, density potential field data, and background structure potential field data separately, and combine them with locally stored object feature data for precise modulation and rendering, ultimately generating the target image of the target object.

[0264] In summary, this embodiment provides a unified spatial carrying plane for the three subfield data by establishing a standardized image spatial grid. Each subfield data is generated, aligned, and combined under the same spatial coordinate system, avoiding spatial misalignment and ensuring spatial consistency during subsequent local processing. Each subfield data is independently encoded and stored, facilitating local modulation of different dimensions and enabling refined visual enhancement control. Furthermore, each subfield data is independent of the precise outline or identity features of the target object, ensuring both scene layout and subject positioning consistency while maintaining privacy and security. Additionally, merging the three subfields into a multi-channel tensor reduces the amount of data transmitted compared to transmitting a complete target object pixel image, decreasing data transmission frequency and protocol overhead, and improving edge-cloud transmission efficiency. Moreover, user preference parameters in the rendering requirements are mapped to the generation process of each subfield data, ensuring that the intermediate representation data generated by the server fully reflects the user's personalized needs and enhancing the user experience.

[0265] It should be noted that the server-side image processing method provided in this embodiment can be implemented independently of the terminal-side method, or it can be used in conjunction with the terminal-side method to form a complete end-to-cloud collaborative system. In an independent implementation scenario, the intermediate representation data generated by the server can be used by any terminal device with object feature data, exhibiting good versatility and flexibility. That is, different terminal devices need to render objects with similar forms, such as different models of mobile phones from the same series, different styles of furniture of the same style, or a series of products from the same brand. The intermediate representation data corresponding to the object with that form generated by the server can be reused, combined with the local object's real object feature data, to render and generate the target image of the object. In this way, the server does not need to execute the above intermediate representation data generation process separately for multiple objects with similar forms, reducing redundant calculations on the server, improving processing efficiency, and saving computing resources and network transmission overhead.

[0266] Based on the above analysis, the server can maintain an intermediate representation data cache. When the server generates intermediate representation data for a target object for the first time, it associates this intermediate representation data with the corresponding reference feature data (especially the target space attribute information and rendering requirement information) and stores it in the cache. Thus, during the generation of intermediate representation data based on the reference feature data, it can first check whether there is already stored intermediate representation data for other objects that match the reference feature data and / or rendering requirement information of the target object. If not, the current intermediate representation data can be generated based on the reference feature data and rendering requirement information according to the method described above; the implementation process is not elaborated in this embodiment. If it exists, the stored intermediate representation data of other objects can be directly obtained as the intermediate representation data of the target object.

[0267] Optionally, in the above detection process, it can be achieved by detecting whether multiple objects (including the target object and at least one other object) have similar target spatial attribute information. That is, the difference between the position region, size range or orientation trend of multiple objects in the image space is less than a preset threshold. For example, if the size ratio difference between the target object and the bounding rectangle of another object does not exceed 10% and the center position offset does not exceed 5% of the image size, it can be considered that the other object has similar target spatial attribute information to the target object, and the intermediate representation data of the other object can be directly reused.

[0268] Optionally, this application can also detect whether the target object is similar in shape and structure to at least one other object. For example, it can determine the similarity of coarse-grained geometric shapes (such as the concavity and convexity features of the overall outline and the relative layout of the main components) after removing visual identity features. If the similarity is higher than the shape similarity threshold, the intermediate representation data of the other object can be reused. In addition, this application can also be compatible with rendering requirements. That is, it can calculate the similarity of the rendering requirement information corresponding to the target object and at least one other object. If the similarity is greater than the rendering similarity threshold, it means that the rendering requirements of these multiple objects are the same or compatible, and the intermediate representation data of the other object can be reused. It can also directly adjust the reused intermediate representation data based on the rendering requirement information of the target object to obtain intermediate representation data adapted to the target object. The adjustment process is not detailed in this application. The threshold for the above similarity determination can be configured or adjusted according to the actual application scenario to balance the reuse rate and the adaptation accuracy.

[0269] In some embodiments, during the adaptive adjustment of reused intermediate representation data based on the difference in target spatial attribute information between the target object and other objects, the activation method described above can be used, or at least one spatial transformation adaptation operation such as translation, scaling, and rotation can be used. Translation adaptation refers to translating the data of each subfield as a whole to align the occupants of the two objects. During scaling adaptation, since latent variable / potential field data of different resolutions cannot be directly reused, interpolation scaling is required. During rotation adaptation, the spatial orientation vector of the illumination potential field and the spatial guidance distribution of the background structural potential field of the two objects need to be rotated synchronously; the implementation method is not detailed in this application.

[0270] In addition, based on the differences in rendering requirements between the target object and other objects, the adaptation process of the reused intermediate representation data can achieve adjustments such as lighting direction, color atmosphere, decoration density, and background type. Specifically, it can offset and adjust the light direction channel of the lighting potential field data in the reused intermediate representation data, convert the color temperature of the color channel, and globally scale the values ​​of each channel of the density potential field data, and replace or mix the geometric construction channel of the background structure potential field data.

[0271] As can be seen, when multiple objects with similar forms have the same or similar reference feature data, the server does not need to repeatedly perform the intermediate representation data generation calculation for each object, significantly saving server computing resources. It can directly read the stored intermediate representation data from the cache, resulting in a much faster response time than regeneration, thus improving the user experience. When encountering mismatched reference feature data, the server automatically generates new intermediate representation data and stores it in the cache. As the cache library is continuously enriched, the probability of subsequent reuse gradually increases.

[0272] Reference Figure 10 This is a schematic diagram of the structure of an image processing device according to Embodiment 1 of this application, applied to a terminal device, such as... Figure 10 As shown, the image processing apparatus may include:

[0273] The object feature data determination module 101 is used to determine the object feature data of the target object; the object feature data includes the visual identity feature information of the target object.

[0274] The reference feature data generation module 102 is used to generate reference feature data of the target object based on the object feature data. The reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can restore the visual identity of the target object.

[0275] Reference feature data sending module 103 is used to send the reference feature data to the server;

[0276] The intermediate representation data receiving module 104 is used to receive intermediate representation data generated by the server based on the reference feature data, wherein the intermediate representation data is used to characterize the distribution trend of visual elements in the image space;

[0277] The target image generation module 105 is used to generate a target image of the target object based on the object feature data and the intermediate representation data.

[0278] Optionally, the reference feature data sending module 103 may include:

[0279] The extraction unit is used to extract the original spatial attribute information of the target object in the image space from the object feature data;

[0280] An information processing unit is used to remove visual feature information that can restore visual identity features from the original spatial attribute information and to perform dimensionality reduction processing on the information to obtain the target spatial attribute information of the target object; the spatial complexity of the target spatial attribute information is lower than the spatial complexity of the original spatial attribute information.

[0281] The determining unit is used to determine the target spatial attribute information as at least a portion of the reference feature data to be transmitted.

[0282] Optionally, the target image generation module 105 may include:

[0283] The subfield data processing unit is used to process the corresponding subfield data in the intermediate representation data based on at least one visual identity feature information in the object feature data to obtain enhanced intermediate representation data that matches the target object.

[0284] A rendering generation unit is used to render and generate a target image of the target object based on the enhanced intermediate representation data;

[0285] Optionally, the image processing apparatus may further include:

[0286] An adjustment module is used to adjust the intermediate representation data in response to interactive operations on the target object in the image space, so as to obtain adjusted intermediate representation data.

[0287] Optionally, the subfield data processing unit may include:

[0288] The first adjustment subunit is used to adjust the spatial distribution of the illumination potential field data or density potential field data in the intermediate representation data within the corresponding contour area based on the boundary features in the object feature data.

[0289] The second adjustment subunit is used to adjust the numerical intensity of the component corresponding to the contact area in the illuminance field data in the intermediate representation data based on the contact area features in the object feature data.

[0290] The third adjustment subunit is used to adjust the optical parameters corresponding to the illumination potential field data and / or density potential field data in the intermediate representation data based on the material property features in the object feature data;

[0291] The fourth adjustment subunit is used to adjust the spatial guidance distribution corresponding to the background structure potential field data in the intermediate representation data or the illumination gradient direction corresponding to the illumination potential field data based on the visual centroid features in the object feature data.

[0292] The fifth adjustment subunit is used to adjust the response intensity of the illumination potential field data, density potential field data and / or background structure potential field data in the intermediate representation data within the protected area based on the protected area characteristics of the target object in the object feature data.

[0293] Reference Figure 11 This is a schematic diagram of the structure of an image processing device according to Embodiment 2 of this application, applied to a server, such as... Figure 11As shown, the image processing apparatus may include:

[0294] The reference feature data receiving module 111 is used to receive reference feature data of a target object sent by a terminal device. The reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can be used to reconstruct the visual identity of the target object.

[0295] The intermediate representation data generation module 112 is used to generate intermediate representation data based on the reference feature data, wherein the intermediate representation data is used to characterize the distribution trend of visual elements in the image space;

[0296] The intermediate representation data sending module 113 is used to send the intermediate representation data to the terminal device so that the terminal device generates a target image of the target object based on the object feature data of the target object and the intermediate representation data, wherein the object feature data includes the visual identity feature information of the target object.

[0297] Optionally, the image processing apparatus may further include:

[0298] A rendering requirement information receiving module is used to receive rendering requirement information for the target object sent by the terminal device;

[0299] Correspondingly, the intermediate representation data generation module 112 may include:

[0300] The first acquisition unit is used to acquire target spatial attribute information from the reference feature data;

[0301] A building unit is used to build a standardized image space grid based on the output resolution in the rendering requirement information;

[0302] The generation unit is used to generate illumination potential field data, density potential field data, and background structure potential field data on the standardized image space grid based on the target spatial attribute information and the rendering requirement information, so as to serve as intermediate representation data adapted to the reference feature data.

[0303] Optionally, the intermediate representation data generation module 112 may further include:

[0304] The detection unit is used to detect whether there is any stored intermediate representation data of other objects that match the reference feature data and / or rendering requirement information of the target object; if not, it triggers the generation unit to generate the current intermediate representation data based on the reference feature data and the rendering requirement information.

[0305] The second acquisition unit is used to acquire, if present, the intermediate representation data of the other objects that have been stored as the intermediate representation data of the target object.

[0306] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, they enable a terminal device or server to implement any of the image processing methods provided in this application.

[0307] The computer-readable storage medium can be any available medium that an electronic device can store, or a data storage device such as a training device or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0308] This application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device enables the electronic device to implement any image processing method executed on the corresponding terminal side or server side provided in this application embodiment.

[0309] When computer-readable instructions are loaded and executed on an electronic device, all or part of the processes or functions described in the embodiments of this application are generated. The electronic device may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer-readable instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer-readable instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means, depending on the actual application scenario.

[0310] This application also provides an intelligent program (such as an agent or intelligent assistant) to receive target object and its rendering requirement information, and implement the image processing method proposed in this application that is executed on the terminal device side. The implementation process can be referred to the description of the corresponding part of the method embodiment applied to the terminal side above. In this implementation process, other components of the application program or operating system can also be controlled through interface calls or other interactive methods to respond to the image processing request of the target object, execute corresponding tasks, and output the target image of the target object. The implementation process is not described in detail in this application. Optionally, if the intelligent program is deployed on the server side, it can receive reference feature data sent by the terminal side and implement the image processing method proposed in this application that is executed on the server side. The implementation process is not described in detail in this application.

[0311] This application also proposes an electronic device that may include at least one memory and at least one processor. When the electronic device is a terminal device, the memory stores a first computer program that implements the image processing method executed on the terminal side, and the processor loads and executes the first computer program to implement the image processing method on the terminal side. When the electronic device is a server, the memory stores a second computer program that implements the image processing method executed on the server side, and the processor loads and executes the second computer program to implement the image processing method on the server side. The implementation process is not described in detail in this application.

[0312] In some embodiments, when the electronic device is a terminal device, its memory can store the program of the intelligent agent, and the processor runs the processor's program to execute the image processing method described above for the terminal-side embodiment. Alternatively, the processor can run the intelligent agent, which executes a first computer program through the processor to implement the steps of the image processing method of the terminal-side embodiment. In this case, the intelligent agent can be configured to implement the terminal-side image processing method of this application, and the implementation process will not be described in detail in this application.

[0313] It should be understood that the structure of the electronic device does not constitute a limitation on the electronic device in the embodiments of this application. In practical applications, the electronic device may include more or fewer components than described above, or combine certain components. When the electronic device is a terminal device, the electronic device may also include sensing units such as gyroscopes, accelerometers and gravity sensors for obtaining sensing parameters, power management modules, antennas or other communication elements, etc. This application will not provide detailed examples of each of these.

[0314] Finally, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0315] In the above embodiments, the invention can be implemented entirely or partially by software, hardware, firmware, or any combination thereof. Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware, or it can be implemented using dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memory, dedicated components, etc. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The various embodiments in this specification are described in a progressive or parallel manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to mutually. For the apparatus, electronic device, and system disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

Claims

1. An image processing method, comprising: Determine the object feature data of the target object; the object feature data includes the visual identity feature information of the target object; Based on the object feature data, reference feature data of the target object is generated. The reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can restore the visual identity of the target object. Send the reference feature data to the server; Receive intermediate representation data generated by the server based on the reference feature data, wherein the intermediate representation data is used to characterize the distribution trend of visual elements in the image space; Based on the object feature data and the intermediate representation data, a target image of the target object is generated.

2. The method according to claim 1, wherein generating reference feature data of the target object based on the object feature data comprises: Extract the original spatial attribute information of the target object in the image space from the object feature data; The original spatial attribute information is subjected to visual feature information that can restore visual identity features and dimensionality reduction processing to obtain the target spatial attribute information of the target object; the spatial complexity of the target spatial attribute information is lower than that of the original spatial attribute information. The target spatial attribute information is determined as at least a portion of the reference feature data to be transmitted.

3. The method according to claim 1, wherein, The intermediate representation data includes visual potential field data, which contains at least one of the following subfield data: Illumination potential field data is used to characterize the distribution of light intensity, chromaticity, or ray parameters in the image space; Density potential field data is used to characterize the distribution density of particles, textures, or lines in image space; Background structure potential field data is used to characterize the distribution of spatial hierarchy, geometric structure, or background topology in image space.

4. The method according to any one of claims 1-3, wherein generating the target image of the target object based on the object feature data and the intermediate representation data comprises: Based on at least one visual identity feature information in the object feature data, the corresponding subfield data in the intermediate representation data is processed to obtain enhanced intermediate representation data that matches the target object; Based on the enhanced intermediate representation data, a target image of the target object is rendered and generated; The visual identity feature information includes at least one of boundary features, contact area features, material property features, visual center of gravity features, and protected area features; The subfield data includes at least one of illumination potential field data, density potential field data, and background structure potential field data.

5. The method according to claim 4, further comprising, before generating the target image of the target object based on the object feature data and the intermediate representation data: In response to interactive operations on the target object in the image space, the intermediate representation data is adjusted to obtain adjusted intermediate representation data; The interactive operation includes adjusting at least one of the position, size, and orientation of the target object in the image space, and / or inputting adjustment prompts for the intermediate representation data.

6. The method according to claim 4, wherein processing the corresponding subfield data in the intermediate representation data based on at least one visual identity feature information in the object feature data includes at least one of the following implementation methods: Based on the boundary features in the object feature data, adjust the spatial distribution of the illumination potential field data or density potential field data in the intermediate representation data within the corresponding contour area; Based on the contact area features in the object feature data, adjust the numerical intensity of the component corresponding to the contact area in the illuminance field data of the intermediate representation data; Based on the material property features in the object feature data, adjust the optical parameters corresponding to the illumination potential field data and / or density potential field data in the intermediate representation data; Based on the visual centroid features in the object feature data, adjust the spatial guidance distribution corresponding to the background structure potential field data in the intermediate representation data or the illumination gradient direction corresponding to the illumination potential field data. Based on the protected area characteristics of the target object in the object feature data, adjust the response intensity of the illumination potential field data, density potential field data and / or background structure potential field data in the intermediate representation data within the protected area.

7. An image processing method, comprising: The terminal device receives reference feature data of a target object, wherein the reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can be used to reconstruct the visual identity of the target object; Intermediate representation data is generated based on the reference feature data, and the intermediate representation data is used to characterize the distribution trend of visual elements in the image space; The intermediate representation data is sent to the terminal device so that the terminal device generates a target image of the target object based on the object feature data of the target object and the intermediate representation data, wherein the object feature data includes the visual identity feature information of the target object.

8. The method according to claim 7, further comprising: Receive rendering requirement information for the target object sent by the terminal device; The generation of intermediate representation data based on the reference feature data includes: Obtain the target spatial attribute information from the reference feature data; Based on the output resolution in the rendering requirement information, a standardized image space grid is established; Based on the target spatial attribute information and the rendering requirement information, illumination potential field data, density potential field data and background structure potential field data are generated on the standardized image spatial grid, respectively, as intermediate representation data adapted to the reference feature data.

9. The method according to claim 8, wherein generating intermediate representation data based on the reference feature data comprises: Detect whether there is stored intermediate representation data of other objects that match the reference feature data and / or rendering requirement information of the target object; If it exists, obtain the intermediate representation data of the other objects that have been stored as the intermediate representation data of the target object; If it does not exist, generate the current intermediate representation data based on the reference feature data and the rendering requirement information.

10. An electronic device comprising at least one memory and at least one processor, wherein: The memory is used to store the program of the intelligent agent; The processor is used to run the agent's program to execute: Determine the object feature data of the target object; the object feature data includes the visual identity feature information of the target object; Based on the object feature data, reference feature data of the target object is generated. The reference feature data is used to characterize the target spatial attribute information of the target object in the image space, and does not contain visual feature information that can restore the visual identity of the target object. Send the reference feature data to the server; Receive intermediate representation data generated by the server based on the reference feature data, wherein the intermediate representation data is used to characterize the distribution trend of visual elements in the image space; Based on the object feature data and the intermediate representation data, a target image of the target object is generated.