Aluminum alloy pipe target forging forming method and system based on image processing
By establishing a three-dimensional digital model of the aluminum alloy tube target and image processing technology, combined with a dual-spectral camera array and a neural radiation field network, accurate monitoring and parameter optimization of the aluminum alloy tube target forging process are achieved, solving the problems of insufficient precision and unstable quality in the existing technology and improving production efficiency and consistency.
Patent Information
- Application Number
- CN202511166468.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-20
AI Technical Summary
The existing technology has problems such as insufficient precision, unstable quality, low efficiency and poor consistency in the forging process of aluminum alloy tube targets, which makes it difficult to meet the needs of modern industry.
An image processing-based method is used to establish a three-dimensional digital model of the aluminum alloy tube target, combine a dual-spectral camera array and a neural radiation field network for image acquisition and reconstruction, and use a multi-scale texture decoupling network and graph structured computing to achieve precise monitoring of the forging process and parameter optimization.
It significantly improves the forging quality and consistency of aluminum alloy tube targets, reduces scrap rate, shortens production cycle, reduces manual intervention, and has strong adaptability and scalability.
Smart Images

Figure CN120655867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology, and in particular to a forging method and system for an aluminum alloy tube target based on image processing. Background Art
[0002] Aluminum alloy tubular targets, key components in advanced manufacturing, have long faced technical bottlenecks in their forging process, including insufficient precision and inconsistent quality. Traditional forging methods, which rely primarily on empirical parameter settings and manual adjustments, are ill-suited to the high-precision forming requirements of complex aluminum alloy tubular targets. This approach is not only inefficient but also highly dependent on the operator's skill level, resulting in poor product consistency and high scrap rates, making it unable to meet the demands of modern industry for high-quality aluminum alloy tubular targets.
[0003] With the development of industrial automation and intelligent manufacturing, image processing technology has been widely used in the manufacturing field. However, its application in the forging process of aluminum alloy tube targets still faces many limitations. Existing technologies mostly use single-spectrum image acquisition, which cannot effectively capture the complex optical properties of aluminum alloys during high-temperature forging. Furthermore, the processing methods for the acquired image data are relatively simple and lack a deep understanding of the material deformation mechanism. This leads to inaccurate forging parameter optimization and difficulty in predicting and controlling the deformation trends of aluminum alloy tube targets during the forging process. Summary of the Invention
[0004] In view of the deficiencies in the prior art, the present invention provides a forging method and system for an aluminum alloy tube target based on image processing, which can solve the problems in the prior art.
[0005] A first aspect of an embodiment of the present invention provides a forging method for an aluminum alloy tube target based on image processing, comprising: Establishing a three-dimensional digital model of the aluminum alloy tube target, dividing the three-dimensional digital model into grid units, and establishing a topological relationship diagram of the grid units; A dual-spectral camera array is used to capture forged images of an aluminum alloy tube target from multiple perspectives, and adaptive exposure compensation is performed to obtain a multi-perspective fusion image. Camera pose parameters are extracted from the multi-perspective fusion image to obtain corresponding viewing direction information. Sine position encoding is performed on the three-dimensional spatial coordinates of the grid cells and the viewing direction information, respectively, and the results are input into a neural radiation field network for volume rendering reconstruction to obtain a temperature radiation intensity mapping image. Phase consistency registration and wavelet transform denoising are performed on the temperature radiation intensity mapping image to obtain an enhanced aluminum alloy tube target image; the enhanced aluminum alloy tube target image is input into a multi-scale texture decoupling network, a local texture descriptor is obtained through multi-scale decomposition and texture feature extraction, the local texture descriptor is input into a component separation unit, a reflective component, an oxidation component, and a forging component are decoupled and separated, and contour feature data of the aluminum alloy tube target is output; Based on the topological relationship diagram of the grid unit and the contour feature data, the deformation trend is predicted through graph structured calculation to generate a forging parameter adjustment instruction; and the forging equipment is controlled according to the forging parameter adjustment instruction.
[0006] In an optional embodiment: The steps of establishing a three-dimensional digital model of an aluminum alloy tube target, dividing the three-dimensional digital model into grid units, and establishing a topological relationship diagram of the grid units include: Establishing a three-dimensional NURBS surface model of an aluminum alloy tube target, wherein the three-dimensional NURBS surface model includes a cylindrical surface and an end surface, and the cylindrical surface and the end surface are constructed by control point coordinates, weight factors, and B-spline basis functions; Meshing the three-dimensional NURBS surface model, establishing an adaptive mesh density control function according to the curvature function, wherein the adaptive mesh density control function calculates a mesh unit size based on a minimum mesh size, a maximum mesh size, a directional weight coefficient, and a mesh density control parameter to generate tetrahedral mesh units; A topological relationship graph is constructed based on the tetrahedral grid unit, wherein the topological relationship graph includes a grid node set and a node connection edge set, wherein the grid node set is used as a vertex and the node connection edge set is used as an edge.
[0007] In an optional embodiment: The method comprises the following steps: using a dual-spectral camera array to collect forging images of an aluminum alloy tube target from multiple viewing angles, performing adaptive exposure compensation on the forging images to obtain a multi-view fusion image; extracting camera pose parameters from the multi-view fusion image to obtain corresponding viewing direction information; performing sinusoidal position encoding on the three-dimensional spatial coordinates of the grid unit and the viewing direction information, respectively, and inputting the encoded features into a neural radiation field network; and numerically integrating the neural radiation field network using a volume rendering method to obtain a temperature radiation intensity mapping image. A dual-spectral camera array is used to collect visible light images and near-infrared images of the aluminum alloy tube target, and an adaptive weight function is established based on the scene brightness value to fuse the visible light image and the near-infrared image to obtain a multi-view fused image; Obtaining an intrinsic parameter matrix and an extrinsic parameter matrix of the dual-spectral camera array, calculating a camera optical center position based on the intrinsic parameter matrix and the extrinsic parameter matrix, constructing a viewing angle direction vector from the camera optical center position and an observation point position, and establishing a corresponding relationship between the multi-view fusion image and the viewing angle direction vector; Performing multi-layer sinusoidal position encoding on the three-dimensional spatial coordinates of the aluminum alloy tube target and the viewing angle direction vector, respectively, and inputting the encoded spatial coordinate features and viewing angle features into a neural radiation field network; The neural radiation field network is trained based on the multi-view fusion image and the corresponding view direction vector; the neural radiation field network outputs a volume density parameter and a direction-related radiation intensity parameter, numerically integrates the volume density parameter and the radiation intensity parameter along the direction of the light, calculates the cumulative transmittance, and obtains a continuous radiation field expression; The light cumulative radiation value is calculated based on the continuous radiation field expression, and the light cumulative radiation value is projected and mapped according to the viewing angle direction to obtain a temperature radiation intensity mapping image.
[0008] In an optional embodiment: The step of training the neural radiation field network based on the multi-view fusion image and the corresponding view direction vector includes: Extracting and fusing features of the multi-view fusion image and the view direction vector to obtain a view perception feature; Constructing an adaptive sampling density function in three-dimensional space, wherein the adaptive sampling density function determines the sampling point density according to the angle between the surface normal vector of the sampling point and the viewing direction vector, and performing layered sampling along the ray direction based on the sampling point density to obtain a sampling point sequence; Input the view-aware features into a feature encoder, construct a multi-layer feature pyramid structure, and obtain a multi-scale feature representation through feature transformation and upsampling fusion; extract view-related features from the multi-scale feature representation and establish a feature-guided attention map; Inputting the sampling point sequence and the attention map into the neural radiance field network, and adaptively weighting the sampling point features based on the attention map; A composite loss function based on material properties is constructed, wherein the composite loss function includes a reconstruction loss term and a physical constraint loss term. The physical constraint loss term includes a thermal radiation continuity constraint and a surface oxide layer reflection characteristic constraint. The thermal radiation continuity constraint imposes a physical consistency constraint on the predicted radiation intensity distribution based on Planck's blackbody radiation law. The surface oxide layer reflection characteristic constraint constrains the reflection behavior under different viewing angles based on the optical properties of the high-temperature oxide film of aluminum alloy. The parameters of the neural radiation field network are optimized and trained based on the composite loss function.
[0009] In an optional embodiment: The enhanced aluminum alloy tube target image is input into a multi-scale texture decoupling network, a local texture descriptor is obtained through multi-scale decomposition and texture feature extraction, the local texture descriptor is input into a component separation unit, and the reflective component, the oxidation component, and the forging component are decoupled and separated to output the contour feature data of the aluminum alloy tube target. The steps include: The enhanced aluminum alloy tube target image is down-sampled to generate images at multiple scale levels; local texture features are extracted from the images at the multiple scale levels, nonlinear feature transformation values are calculated within a local neighborhood of a preset size, and the nonlinear feature transformation values are weightedly combined with local weights to obtain a local texture descriptor for each scale level; Inputting the local texture descriptor into a component separation unit, the component separation unit realizes decoupling and separation of the reflective component, the oxidation component, and the forging component through causal reasoning and hierarchical decoupling, thereby obtaining decoupling features at each scale level; Calculating the feature discriminability of the decoupled features at each scale level, determining the corresponding fusion weight coefficient according to the feature discriminability, and performing weighted combination of the decoupled features using the fusion weight coefficient to obtain a fusion feature; An aluminum alloy tube target image is reconstructed based on the fusion feature, and a difference value between the reconstructed image and the original image and a spatial gradient value of the decoupling feature are calculated. The difference value and the spatial gradient value are used as constraints to optimize the parameters of the component separation unit, and the optimized aluminum alloy tube target contour feature data is output.
[0010] In an optional embodiment: The component separation unit realizes the decoupling and separation of the reflective component, the oxidation component, and the forging component through causal reasoning and hierarchical decoupling. The steps of obtaining the decoupling characteristics of each scale level include: The component separation unit includes a causal reasoning module and a hierarchical decoupling module; in the causal reasoning module, a causal graph structure is constructed between the reflective component, the oxidation component, and the forging component, and the conditional probability distribution between the components is calculated based on the causal graph structure to establish the mutual dependence relationship between the component features; in the hierarchical decoupling module, a decoupling loss function including a reliability assessment item is constructed, and the reliability assessment item calculates the decoupling confidence based on the feature distribution of the local area, and adopts an adaptive decoupling strength for different confidence areas; Based on the interdependence and the decoupling confidence, a progressive training strategy is adopted to perform feature decoupling, a decoupling benchmark is established, and the decoupling process is gradually expanded to achieve reliable separation of the reflective component, the oxidation component, and the forging component; The feature reconstruction loss is calculated according to the decoupling benchmark, and the parameters of the causal reasoning module are optimized based on the feature reconstruction loss to obtain the decoupling features of each scale level.
[0011] In an optional embodiment: The steps of predicting deformation trends through graph structured calculation based on the topological relationship graph of the grid units and the contour feature data and generating forging parameter adjustment instructions include: Mapping the contour feature data to a topological relationship graph of the grid cells to construct a graph feature matrix, wherein the graph feature matrix includes geometric features and material state parameters of each grid node; A graph neural network is used to propagate features of the graph feature matrix. A message passing mechanism is established based on the connection relationship between grid nodes. The stress distribution and deformation prediction value of each node are calculated through message aggregation. A deformation state vector is constructed based on the stress distribution and deformation prediction values. The deformation state vector includes the node displacement field, stress field, and material flow characteristics. Comparing the deformation state vector with a preset target shape, calculating the shape deviation, and predicting subsequent deformation trends based on the shape deviation and material deformation law; According to the predicted deformation trend, a forging parameter mapping function is established, wherein the forging parameter mapping function converts the predicted deviation into adjustment amounts of forging force, forging speed and temperature control parameters, and generates a forging parameter adjustment instruction.
[0012] In a second aspect, a forging system for an aluminum alloy tube target based on image processing is provided, comprising: The first unit is used to establish a three-dimensional digital model of the aluminum alloy tube target, divide the three-dimensional digital model into grid units, and establish a topological relationship diagram of the grid units; The second unit is used to use a dual-spectrum camera array to collect forging images of the aluminum alloy tube target from multiple perspectives, perform adaptive exposure compensation to obtain a multi-perspective fusion image; extract camera pose parameters from the multi-perspective fusion image to obtain corresponding perspective direction information; perform sinusoidal position encoding on the three-dimensional spatial coordinates of the grid unit and the perspective direction information, respectively, and input them into a neural radiation field network for volume rendering reconstruction to obtain a temperature radiation intensity mapping image; The third unit is configured to perform phase consistency registration and wavelet transform denoising on the temperature radiation intensity mapping image to obtain an enhanced aluminum alloy tube target image; input the enhanced aluminum alloy tube target image into a multi-scale texture decoupling network, obtain a local texture descriptor through multi-scale decomposition and texture feature extraction, input the local texture descriptor into a component separation unit, perform decoupling and separation of the reflection component, the oxidation component, and the forging component, and output the contour feature data of the aluminum alloy tube target; The fourth unit is used to predict the deformation trend through graph structured calculation based on the topological relationship diagram of the grid unit and the contour feature data, and generate a forging parameter adjustment instruction; and control the forging equipment according to the forging parameter adjustment instruction.
[0013] According to a third aspect, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0014] This invention combines a gridded three-dimensional digital model of an aluminum alloy tube target with dual-spectral image fusion technology to achieve comprehensive monitoring and precise modeling of the forging process. Volume rendering reconstruction using a neural radiation field network accurately captures the temperature radiation characteristics of aluminum alloys at high temperatures, providing high-quality data for subsequent deformation analysis. The innovative application of a multi-scale texture decoupling network and a component separation unit effectively distinguishes between various influencing factors, such as reflection, oxidation, and forging, significantly improving the accuracy and reliability of contour feature extraction.
[0015] This method, based on a graph-structured calculation-based deformation trend prediction method, establishes a mapping relationship between material properties and forming parameters, enabling intelligent optimization and real-time adjustment of forging parameters. This method significantly improves the forging quality and consistency of aluminum alloy tube targets, reduces scrap rates, minimizes manual intervention, and shortens production cycles. Furthermore, this method exhibits strong adaptability and scalability, enabling it to meet the forming requirements of aluminum alloy tube targets of varying specifications and shapes, providing a new technical approach and solution for the field of aluminum alloy precision forging. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 The figure is a flow chart of a forging method of an aluminum alloy tube target based on image processing according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present invention will be described below in conjunction with the drawings in the embodiments of the present invention. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0018] Figure 1 This is a flow chart of a forging method for an aluminum alloy tube target based on image processing according to the present invention, as shown in FIG. Figure 1 As shown, the method includes: Establishing a three-dimensional digital model of the aluminum alloy tube target, dividing the three-dimensional digital model into grid units, and establishing a topological relationship diagram of the grid units; A dual-spectral camera array is used to capture forged images of an aluminum alloy tube target from multiple perspectives, and adaptive exposure compensation is performed to obtain a multi-perspective fusion image. Camera pose parameters are extracted from the multi-perspective fusion image to obtain corresponding viewing direction information. Sine position encoding is performed on the three-dimensional spatial coordinates of the grid cells and the viewing direction information, respectively, and the results are input into a neural radiation field network for volume rendering reconstruction to obtain a temperature radiation intensity mapping image. Phase consistency registration and wavelet transform denoising are performed on the temperature radiation intensity mapping image to obtain an enhanced aluminum alloy tube target image; a phase consistency registration method is used to align multiple frames of images, phase information of the images is extracted through Fourier transform, and the displacement relationship between images is determined using the phase correlation maximization principle to achieve sub-pixel level precise registration; wavelet transform denoising technology is applied to decompose the image into different frequency sub-bands, and a soft threshold or hard threshold shrinkage method is used for each sub-band coefficient to suppress the noise component, and an enhanced aluminum alloy tube target image is reconstructed through an inverse wavelet transform, retaining edge details while effectively suppressing noise interference generated during the thermal imaging process; Inputting the enhanced aluminum alloy tube target image into a multi-scale texture decoupling network, obtaining a local texture descriptor through multi-scale decomposition and texture feature extraction, inputting the local texture descriptor into a component separation unit, performing decoupling and separation of the reflective component, the oxidation component, and the forging component, and outputting the contour feature data of the aluminum alloy tube target; Based on the topological relationship diagram of the grid units and the contour feature data, the deformation trend is predicted through graph structured calculation to generate forging parameter adjustment instructions; the forging equipment is controlled according to the forging parameter adjustment instructions, and optimization and adjustment are performed based on real-time feedback data until the forming quality of the aluminum alloy tube target reaches a preset standard.
[0019] In an optional embodiment: The steps of establishing a three-dimensional digital model of an aluminum alloy tube target, dividing the three-dimensional digital model into grid units, and establishing a topological relationship diagram of the grid units include: Establishing a three-dimensional NURBS surface model of an aluminum alloy tube target, wherein the three-dimensional NURBS surface model includes a cylindrical surface and an end surface, and the cylindrical surface and the end surface are constructed by control point coordinates, weight factors, and B-spline basis functions; Meshing the three-dimensional NURBS surface model, establishing an adaptive mesh density control function according to the curvature function, wherein the adaptive mesh density control function calculates a mesh unit size based on a minimum mesh size, a maximum mesh size, a directional weight coefficient, and a mesh density control parameter to generate tetrahedral mesh units; A topological relationship graph is constructed based on the tetrahedral grid unit, wherein the topological relationship graph includes a grid node set and a node connection edge set, wherein the grid node set is used as a vertex and the node connection edge set is used as an edge.
[0020] For example, a three-dimensional NURBS surface model of an aluminum alloy tube target is established. The aluminum alloy tube target mainly consists of two parts: a cylindrical surface and an end face. When constructing the cylindrical surface, the dimensional parameters of a radius of 15 mm and a length of 120 mm are selected. 12 control points are set along the circumferential direction and 10 control points are set along the axial direction to form a control point grid. Each control point is assigned a three-dimensional coordinate value and a corresponding weight factor. For example, the weight factor of the first row of control points selected on the cylindrical surface is set to 0.7, and the weight factor of the control points in the remaining rows is set to 1.0 to ensure accurate expression of the surface shape. When constructing the end face, 8 control points are used to form a closed circular end face. The control points are distributed on a circle with a radius of 15 mm, and the weight factors are all set to 0.9. For the order selection of the NURBS basis function, a 3rd-order B-spline basis function is used in the circumferential direction to accurately express the circular cross-section; a 2nd-order B-spline basis function is used in the axial direction to meet the shape change requirements of the tube target along the axial direction. The node vectors are uniformly distributed, with the node vectors in the circumferential direction being [0, 0, 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1, 1] and in the axial direction being [0, 0, 0.125, 0.25, 0.375, 0.5, 0.625, 0.75, 0.875, 1, 1]. The precise construction of the three-dimensional NURBS surface model of the aluminum alloy tube target is achieved through the combination of control point coordinates, weight factors, basis function orders, and node vectors.
[0021] Because different parts of the aluminum alloy tube target may experience varying degrees of deformation during the forging process, an adaptive mesh density control function is required based on the geometric characteristics. First, the curvature values of each point on the 3D NURBS surface model are calculated. For the cylindrical portion of the tube target, a principal curvature radius of 15 mm corresponds to a curvature value of 0.067. In the transition region of the end face, the curvature value increases significantly, reaching a maximum of 0.2. Based on the curvature distribution, an adaptive mesh density control function is established. The minimum mesh size is set to 0.5 mm, and the maximum mesh size is set to 3 mm. In areas with greater curvature, such as the transition region between the end face and the cylindrical surface, the mesh size approaches the minimum value of 0.5 mm. In areas with less curvature, such as the central region of the cylindrical surface, the mesh size approaches the maximum value of 3 mm. A directional weighting factor is also introduced to achieve different mesh densities in the axial and circumferential directions. The axial weighting factor is set to 1.2, and the circumferential weighting factor is set to 0.9 to accommodate the deformation characteristics of the tube target in different directions. The mesh density control parameter is set to 0.8 to adjust the mapping between curvature and mesh size. The control function takes the form of a target mesh size that is inversely proportional to the curvature, while also taking into account the influence of directional weights. For example, in an area with a curvature value of 0.1, the axial mesh size is approximately 1.5 mm, and the circumferential mesh size is approximately 2 mm after accounting for the weight coefficients.
[0022] During mesh generation, a triangular mesh is first generated on the model surface, with a control unit count of between 10,000 and 15,000. This mesh is then expanded inward to generate tetrahedral mesh elements, totaling approximately 30,000 to 40,000 elements. The generated mesh must meet quality standards: a minimum dihedral angle greater than 15 degrees, a maximum dihedral angle less than 160 degrees, and a volume ratio greater than 0.2 to ensure numerical stability. A topological graph is constructed based on the tetrahedral mesh elements. Each mesh node is treated as a vertex of the graph, and the connecting edges between nodes are treated as edges. For a typical aluminum alloy tube target model, the number of mesh nodes is approximately 8,000 to 10,000, and the number of connecting edges is approximately 45,000 to 55,000. Each mesh node stores its spatial coordinate information (x, y, z), and each connecting edge records the index of the two nodes it connects.
[0023] In the topology diagram, each node is assigned an attribute label, including information such as node type (surface node or internal node) and location (cylindrical region or end face region). For surface nodes, surface normal information is additionally stored for subsequent deformation analysis. Furthermore, each edge is assigned a weight, proportional to its length, which is used for weight calculation in subsequent mesh deformation analysis.
[0024] Furthermore, an adjacency matrix or adjacency table is constructed within the topological graph to record the connectivity between nodes. For a grid with n nodes, the adjacency matrix is an n×n sparse matrix, where nonzero elements indicate a connection between nodes. Given the sparse nature of aluminum alloy tube target grids, an adjacency list representation is generally used, with an average connectivity degree of approximately 5 to 7 per node.
[0025] This method fully considers the geometric properties and deformation characteristics of the aluminum alloy tube target, accurately expresses complex geometric shapes through NURBS surfaces, ensures the balance between calculation accuracy and efficiency through adaptive meshing, and captures the spatial relationship between mesh units through topological relationship diagrams.
[0026] In an optional embodiment: The method comprises the following steps: using a dual-spectral camera array to collect forging images of an aluminum alloy tube target from multiple viewing angles, performing adaptive exposure compensation on the forging images to obtain a multi-view fusion image; extracting camera pose parameters from the multi-view fusion image to obtain corresponding viewing direction information; performing sinusoidal position encoding on the three-dimensional spatial coordinates of the grid unit and the viewing direction information, respectively, and inputting the encoded features into a neural radiation field network; and numerically integrating the neural radiation field network using a volume rendering method to obtain a temperature radiation intensity mapping image. Evenly arranging multiple observation positions in a spherical coordinate system, using a dual-spectral camera array to collect visible light images and near-infrared images of the aluminum alloy tube target at the observation positions, establishing an adaptive weight function based on scene brightness values, and fusing the visible light image and the near-infrared image according to the adaptive weight function to obtain a multi-view fused image; A calibration plate is used to obtain an intrinsic parameter matrix and an extrinsic parameter matrix of the dual-spectral camera array, the camera optical center position is calculated based on the intrinsic parameter matrix and the extrinsic parameter matrix, a viewing angle direction vector is constructed by combining the camera optical center position and the observation point position, and a corresponding relationship between the multi-view fusion image and the viewing angle direction vector is established; Multi-layer sinusoidal position coding is performed on the three-dimensional spatial coordinates of the aluminum alloy tube target and the viewing direction vector, respectively. The multi-layer sinusoidal position coding maps the input features to a high-dimensional space through sine-cosine functions of different frequencies, and the encoded spatial coordinate features and viewing direction features are input into a neural radiation field network; The neural radiation field network is trained based on the multi-view fusion image and the corresponding view direction vector; the neural radiation field network outputs a volume density parameter and a direction-related radiation intensity parameter, numerically integrates the volume density parameter and the radiation intensity parameter along the direction of the light, calculates the cumulative transmittance, and obtains a continuous radiation field expression; The light cumulative radiation value is calculated based on the continuous radiation field expression, and the light cumulative radiation value is projected and mapped according to the viewing angle direction to obtain a temperature radiation intensity mapping image.
[0027] For example, to ensure full coverage of the aluminum alloy tube target, 12 observation points are set up on a hemisphere with a radius of 1.5 meters, centered on the tube target. These observation points are divided into three layers in the zenith angle direction, located at 30 degrees, 60 degrees, and 75 degrees respectively. In azimuth, they are evenly distributed, with four points per layer and an azimuth interval of 90 degrees between adjacent points. This arrangement ensures full observation of the tube target surface while avoiding blind spots. A dual-spectral camera is installed at each observation location. The camera consists of two sensors: a visible light sensor (400-700nm) and a near-infrared sensor (700-1100nm). The two sensors separate their optical paths via a beam splitter prism, ensuring simultaneous acquisition of the same scene. At each observation location, both visible light and near-infrared images are collected simultaneously. During the forging process, the temperature of the aluminum alloy tube target typically ranges from 400-600°C, with predominantly red-orange radiation in the visible spectrum and a stronger thermal radiation signal in the near-infrared spectrum. Since visible light images are prone to overexposure at high temperatures, and near-infrared images are more sensitive to temperature changes, it is necessary to fuse the two images to obtain more comprehensive information.
[0028] Before image fusion, adaptive exposure compensation is performed. For visible light images, the proportion of overexposed areas is calculated through histogram analysis. When the overexposure ratio exceeds 10%, the exposure time is adjusted to 70% of the original value. For near-infrared images, when the image average grayscale value falls below 30% of the total grayscale range, the exposure time is increased to 120% of the original value. This ensures optimal image quality under different temperature conditions.
[0029] Image fusion is performed by establishing an adaptive weighting function based on scene brightness. For each pixel, the normalized brightness values of the visible and near-infrared images are calculated. When the visible light pixel value exceeds 240 (8-bit image, maximum value 255), the visible light weight at that location is set to 0.3, and the near-infrared weight is set to 0.7. When the visible light pixel value is below 50, the visible light weight at that location is set to 0.6, and the near-infrared weight is set to 0.4. For intermediate brightness regions, the weights vary linearly with brightness. Based on this weighting strategy, the two images are weightedly fused to produce a multi-view fused image.
[0030] A standard checkerboard calibration plate is used for camera calibration. The calibration plate is in a 9×7 format, with each grid measuring 25mm×25mm. During the calibration process, the calibration plate is placed at 12 different positions and angles to ensure that the entire field of view of the camera is covered. By identifying the corner points of the calibration plate, the camera's intrinsic parameter matrix and distortion coefficients are calculated. The intrinsic parameter matrix contains the focal length and principal point coordinates. For example, in a typical intrinsic parameter matrix, the focal length value is approximately 1200 pixels, and the principal point coordinates are located near the center of the image. The distortion coefficient includes radial distortion and tangential distortion parameters. Typically, the radial distortion parameter k1 is approximately -0.1, and k2 is approximately 0.02. The extrinsic parameter matrix is calculated by setting the origin of the world coordinate system at a fixed reference point on the forging equipment, using the known camera installation position and orientation. For example, for a camera located at a zenith angle of 30 degrees and an azimuth angle of 0 degrees, the rotation part of its extrinsic parameter matrix represents the rotation angle of the camera coordinate system relative to the world coordinate system, and the translation part represents the position coordinate of the camera optical center in the world coordinate system, which is approximately (1.3, 0, 0.75) meters. Based on the intrinsic parameter matrix and the extrinsic parameter matrix, the optical center position and principal axis direction of each camera are calculated. The optical center position is obtained directly from the translation part of the extrinsic parameter matrix, while the principal axis direction is calculated by the rotation part of the extrinsic parameter matrix. The optical center position is connected to the observation point on the aluminum alloy tube target to construct the viewing direction vector. For a typical setting, the viewing direction vector covers all angles from vertical observation to nearly horizontal observation, ensuring full capture of the tube target surface.
[0031] For three-dimensional spatial coordinates, 10 frequency levels are selected, starting with a base frequency of 20 and increasing by 1.5 per level, forming a frequency sequence: 20, 30, 45, 67.5, and so on. For each frequency, the sine and cosine values of the coordinate in three directions are calculated, resulting in a 60-dimensional encoding vector (10 frequencies × 3 directions × 2 functions). Similarly, the view direction vector is encoded using 8 frequency levels, resulting in a 48-dimensional direction encoding vector (8 frequencies × 3 directions × 2 functions). This high-dimensional encoding method effectively overcomes the difficulty of neural networks in learning low-frequency functions. The encoded spatial coordinate features and view direction features are input into a neural radiance field network. This network uses a multi-layer perceptron architecture with 8 hidden layers, 256 neurons per layer, and a ReLU activation function. The network is divided into two parts: the first five layers encode the spatial coordinates and output volume density values; the last three layers combine the spatial and directional encodings to output RGB color values and radiance intensity values. The training dataset contains tens of thousands of ray samples, with 64 points upsampled per ray. Training was performed using the Adam optimizer, with an initial learning rate of 0.0005, which was halved every 50,000 steps. After training, the neural radiation field network was able to predict the volume density and radiation intensity parameters for any three-dimensional point. To generate the temperature radiation intensity map, the predicted volume density and radiation intensity were numerically integrated along each ray. Specifically, for each ray, 128 points were uniformly sampled from 0.5 meters at the near end to 2 meters at the far end. The volume density and radiation intensity at each sampled point were then calculated, followed by the cumulative transmittance. The cumulative transmittance represents the proportion of unabsorbed light remaining when the ray reaches a certain point, decreasing from far to near. Finally, the cumulative radiation value of the ray was calculated based on the continuous radiation field representation. The volume density value was multiplied by the radiation intensity value, then by the cumulative transmittance, and the final radiation value was obtained by integrating along the ray. These radiation values were projected and mapped according to the viewing angle to form the temperature radiation intensity map. For each point on the target surface, multiple radiation intensity values were obtained from different viewing angles, and the average of these values was taken as the final radiation intensity.
[0032] This method, combining dual-spectral imaging and neural radiation field technology, overcomes the measurement errors caused by traditional thermal imaging due to variations in reflectance and emissivity on high-temperature metal surfaces, significantly improving the accuracy and spatial resolution of temperature field reconstruction. The reconstructed temperature radiation intensity mapping image intuitively displays the surface temperature distribution of the target tube, providing a reliable basis for subsequent adjustments to forging process parameters and effectively improving the forging quality and consistency of aluminum alloy target tubes.
[0033] In an optional embodiment: The step of training the neural radiation field network based on the multi-view fusion image and the corresponding view direction vector includes: Extracting and fusing features of the multi-view fusion image and the view direction vector to obtain a view perception feature; Constructing an adaptive sampling density function in three-dimensional space, wherein the adaptive sampling density function determines the sampling point density according to the angle between the surface normal vector of the sampling point and the viewing direction vector, and performing layered sampling along the ray direction based on the sampling point density to obtain a sampling point sequence; Input the view-aware features into a feature encoder, construct a multi-layer feature pyramid structure, and obtain a multi-scale feature representation through feature transformation and upsampling fusion; extract the view-related features in the multi-scale feature representation, and establish a feature-guided attention map for modulating the feature aggregation process of the sampling points; Inputting the sampling point sequence and the attention map into the neural radiation field network, adaptively weighting the sampling point features based on the attention map, and the neural radiation field network outputting a volume density parameter and a radiation intensity parameter of each sampling point; A composite loss function based on material properties is constructed, wherein the composite loss function includes a reconstruction loss term and a physical constraint loss term. The physical constraint loss term includes a thermal radiation continuity constraint and a surface oxide layer reflection characteristic constraint. The thermal radiation continuity constraint imposes a physical consistency constraint on the predicted radiation intensity distribution based on Planck's blackbody radiation law. The surface oxide layer reflection characteristic constraint constrains the reflection behavior under different viewing angles based on the optical properties of the high-temperature oxide film of aluminum alloy. The parameters of the neural radiation field network are optimized and trained based on the composite loss function.
[0034] For each multi-view fusion image, a ResNet-34 backbone network is used for feature extraction. Feature maps from conv1 to conv5 are retained, with feature dimensions of 64, 128, 256, 512, and 512, respectively. Simultaneously, the view direction vector is encoded using a six-layer fully connected network with 64, 128, 256, 256, 128, and 64 neurons per layer. The activation function is LeakyReLU with a skew coefficient of 0.2. During the feature fusion stage, an attention mechanism is employed to convert the view direction encoding features into channel attention weights. Specifically, the 64-dimensional view encoding features are mapped into a weight vector matching the number of image feature channels through a two-layer fully connected network. Each channel of the image features is then weighted. For example, for the 256-channel feature map output by conv3, a 256-dimensional weight vector is generated, with each weight value ranging from 0.5 to 1.5 to modulate the importance of each channel. This approach allows the network to adaptively adjust feature extraction strategies based on different view directions, forming view-aware features.
[0035] Next, an adaptive sampling density function is constructed in three-dimensional space. To improve sampling efficiency, areas with distinct surface features are prioritized. The cosine of the angle between the surface normal vector at the sampling point and the viewing direction vector is defined as a weighting factor. When the angle is close to 90 degrees (cosine close to 0), the viewing direction is nearly parallel to the surface. These areas often contain more surface detail, so the sampling density should be increased. When the angle is close to 0 or 180 degrees (cosine close to 1), the viewing direction is nearly perpendicular to the surface. These areas have smoother surface features, so the sampling density can be reduced.
[0036] Based on the aforementioned weighting factors, the sampling density function is defined as the product of a base density and a weight modulation term. The base density is set at 80 sampling points per meter, and the weight modulation term ranges from 0.5 to 2.0. In practice, when the absolute value of the weighting factor is less than 0.2, the sampling density is increased to twice the base density; when the absolute value of the weighting factor is greater than 0.8, the sampling density is reduced to 0.5 times the base density; the sampling density in intermediate regions varies linearly. This allows the sampling point spacing to be as small as 0.625 cm in critical areas such as the target's edges and curved transitions, while it can be as large as 2.5 cm in smooth areas such as the cylindrical surface of the target body. For stratified sampling along the ray direction, the near end is set at 0.5 meters and the far end at 2.0 meters. The total number of sampling points is controlled between 64 and 192, dynamically adjusted according to the adaptive sampling density function. For rays with a viewing direction nearly parallel to the target surface, the number of sampling points is close to 192; for rays with a viewing direction nearly perpendicular to the target surface, the number of sampling points is close to 64. Through this adaptive sampling strategy, the sampling accuracy of key areas is guaranteed and the computational complexity is controlled.
[0037] View-aware features are input into the feature encoder to construct a multi-layer feature pyramid. The feature pyramid consists of five scale levels, with the resolution progressively upsampled from 1 / 32 of the original image to the original resolution. At each scale level, features are extracted using a 3×3 convolutional layer with output channels of 512, 256, 128, 64, and 32, respectively. Bilinear interpolation is used for upsampling with an upsampling factor of 2. Feature fusion between adjacent levels is achieved through skip connections. Specifically, the upsampled features of the previous level are concatenated with the features of the current level in the channel dimension, and the number of channels is then adjusted using a 1×1 convolution.
[0038] Viewpoint-dependent features are extracted from the multi-scale feature representation to construct a feature-guided attention map. The attention map generation process consists of two steps: first, the features encoded by the view direction vector are mapped into a weight vector with the same number of channels as the feature map through a fully connected layer; then, for each spatial position in the feature map, the dot product of the feature vector and the weight vector is calculated and normalized to the range of 0 to 1 using the Sigmoid function to form a spatial attention map. In practical applications, the distribution of attention values for different regions of the aluminum alloy tube target, such as the middle of the cylindrical surface, the end face, and the transition region, shows significant differences. Typically, the attention values for edges and transition regions range from 0.7 to 0.9, while the attention values for smooth regions range from 0.3 to 0.5.
[0039] The sampling point sequence and attention map are input into the Neural Radiance Field Network, which adaptively weights the sampling point features based on the attention map. The Neural Radiance Field Network uses an 8-layer fully connected network with 256 neurons per layer and a Reinforced Luminance (ReLU) activation function. The network inputs are the spatial coordinates of the sampling point, the viewing direction vector, and the attention value at the corresponding location; the outputs are the volume density parameter and the radiance intensity parameter. The volume density parameter is a scalar value representing the "probability" of a spatial point, ranging from 0 to 1. A value closer to 1 indicates a higher probability that the point is located on the surface of an object. The radiance intensity parameter is a three-dimensional vector corresponding to the three channels of the RGB color space, representing the radiance intensity of the point when viewed from a specific viewing angle.
[0040] During feature aggregation, the sampling point locations are projected onto a two-dimensional feature map, and the feature vectors at the corresponding locations are extracted. These vectors are then weighted by the attention value. This allows the network to adaptively adjust the contribution of features based on the importance of different regions. For regions with high attention values, the sampling point features have a larger weight and a greater impact on the final prediction results. For regions with low attention values, the sampling point features have a smaller weight and a smaller impact on the final prediction results.
[0041] This section does not fully describe the composite loss function, lacking specific calculation methods and implementation details. The following is a supplementary and improved description: A composite loss function based on material properties is constructed, consisting of a reconstruction loss and a physical constraint loss. The reconstruction loss uses mean squared error (MSE) to calculate the difference between the predicted RGB color values and the ground-truth RGB color values. Specifically, the squared difference is calculated for each pixel's RGB channels, and then the squared differences across all pixel locations and channels are summed and averaged. Initially, the reconstruction loss weight is set to 1.0. As training progresses, it is gradually reduced to 0.7 using an exponential decay strategy, with a decay rate of 10% every 50,000 steps, to make the network more focused on physical constraints. The physical constraint loss term includes constraints on thermal radiation continuity and the reflectivity of the surface oxide layer. The thermal radiation continuity constraint is based on Planck's blackbody radiation law and requires that the predicted radiation intensity distribution satisfy continuity with temperature. Specifically, pairs of adjacent sampling points are taken along the ray direction. The difference in radiation intensity between each pair is calculated, for each pair of RGB channels, and the L1 norm (absolute value) is applied as the gradient metric. When the absolute value of the gradient exceeds a preset threshold of 0.05, a quadratic penalty is applied to the excess, with a penalty coefficient of 0.1. The surface oxide layer reflection characteristic constraint takes into account the optical properties of the oxide film formed on the surface of aluminum alloy at high temperatures, requiring that the reflection behavior under different viewing angles conform to physical laws. This is achieved by randomly selecting 4-6 observations from different viewing angles for the same 3D spatial point from the dataset and calculating the Pearson correlation coefficient of the predicted radiation intensity between each pair of viewing angles. Under normal circumstances, the correlation coefficient should not be less than 0.7; when the correlation coefficient is lower than this threshold, the penalty term is calculated as the square of the difference between the threshold and the actual correlation coefficient, multiplied by a penalty coefficient of 0.2. The final composite loss function is the weighted sum of the reconstruction loss, the thermal radiation continuity constraint loss, and the surface oxide layer reflection characteristic constraint loss, with the weight ratio being 0.7:0.1:0.2 in the later stages of training.
[0042] The parameters of the neural radiation field network were optimized and trained based on a composite loss function. The Adam optimizer was used, with an initial learning rate of 0.0001, which was reduced to 50% every 50,000 steps. The batch size was set to 4096 rays, and the number of sampling points on each ray was dynamically adjusted based on an adaptive sampling strategy. The training process was divided into two phases: the first phase, which performed 200,000 iterations and focused on optimizing the reconstruction loss; the second phase, which performed 100,000 iterations and optimized both the reconstruction loss and the physical constraint loss. During training, a model checkpoint was saved every 10,000 steps for subsequent evaluation and selection of the optimal model.
[0043] This method combines advanced feature extraction techniques from computer vision with the physical constraints of materials science to effectively address the variable reflectance and strong viewpoint dependence of high-temperature metal surfaces. The trained network not only reconstructs visually realistic temperature radiation fields but also ensures that the reconstructions adhere to physical laws. This provides a reliable tool for precise monitoring and control of the aluminum alloy tube target forging process, significantly improving the accuracy of forging quality assessment and the efficiency of forging parameter optimization.
[0044] In an optional embodiment: The enhanced aluminum alloy tube target image is input into a multi-scale texture decoupling network, a local texture descriptor is obtained through multi-scale decomposition and texture feature extraction, the local texture descriptor is input into a component separation unit, and the reflective component, the oxidation component, and the forging component are decoupled and separated to output the contour feature data of the aluminum alloy tube target. The steps include: The enhanced aluminum alloy tube target image is subjected to a downsampling operation to generate images at multiple scale levels, wherein the downsampling operation performs a scale reduction process on the aluminum alloy tube target image layer by layer; local texture features are extracted from the images at the multiple scale levels, nonlinear feature transformation values are calculated within a local neighborhood of a preset size, and the nonlinear feature transformation values are weightedly combined with local weights to obtain a local texture descriptor for each scale level; Inputting the local texture descriptor into a component separation unit, the component separation unit realizes decoupling and separation of the reflective component, the oxidation component, and the forging component through causal reasoning and hierarchical decoupling, thereby obtaining decoupling features at each scale level; Calculating the feature discriminability of the decoupled features at each scale level, determining the corresponding fusion weight coefficient according to the feature discriminability, and performing weighted combination of the decoupled features using the fusion weight coefficient to obtain a fusion feature; An aluminum alloy tube target image is reconstructed based on the fusion feature, and a difference value between the reconstructed image and the original image and a spatial gradient value of the decoupling feature are calculated. The difference value and the spatial gradient value are used as constraints to optimize the parameters of the component separation unit, and the optimized aluminum alloy tube target contour feature data is output.
[0045] For example, the enhanced aluminum alloy tube target image is downsampled to generate images at multiple scale levels. Multi-scale decomposition is performed using a Gaussian pyramid structure, with a total of five scale levels. Taking the original image resolution of 1024×768 pixels as an example, after downsampling, images of 512×384, 256×192, 128×96, and 64×48 pixels are generated, respectively. Convolution is performed using a 5×5 Gaussian kernel with a standard deviation of 1.2, followed by a factor of 2 downsampling. This pyramid structure captures image features at different scales and adapts to the varying texture features of the aluminum alloy tube target surface. Local texture features are extracted from the images at each scale level. At each scale level, a local descriptor method is used to extract texture features. Specifically, a nonlinear feature transformation is calculated within a local neighborhood of a preset size, centered at each pixel. For the first level (original resolution), the local neighborhood size is set to 9×9 pixels; for the second level, it is 7×7 pixels; and for levels 3 and below, it is 5×5 pixels. Within each local neighborhood, the following features are computed: local histogram of oriented gradients (8 directional bins), local binary pattern (using a circular neighborhood with a radius of 2, with a total of 8 sampling points), and local contrast feature (calculated as the difference between the maximum and minimum values within the neighborhood). These features are processed using nonlinear transformations. For the histogram of oriented gradients, the cumulative values in each direction are normalized to 1; for the local binary pattern, the binary code is converted to decimal values; and for the local contrast feature, the value is normalized using the sigmoid function to range between 0 and 1. These features are then mapped to a 64-dimensional feature space using a fully connected layer to obtain a preliminary texture description.
[0046] Next, local weights are calculated and weighted together. Local weights are determined based on texture complexity, which is assessed through a comprehensive evaluation of local variance and edge strength. Local variance is obtained by calculating the standard deviation of pixel values within a neighborhood; edge strength is obtained by calculating the gradient magnitude using the Sobel operator. These two metrics are normalized and weighted together with weights of 0.4 and 0.6, respectively, to produce a texture complexity score. The complexity score is converted to weights using a softmax function, summing to 1. The weights are then element-wise multiplied with the aforementioned 64-dimensional texture descriptor to produce a weighted local texture descriptor.
[0047] The local texture descriptor is input into the component separation unit to decouple and separate the reflective, oxidized, and forged components. The component separation unit employs a multi-branch encoder-decoder structure with three parallel branches, one for the reflective, one for the oxidized, and one forged components. Each branch consists of three convolutional layers with a 3×3 kernel size, 64 channels, 128 channels, and 64 channels, respectively. The activation function is LeakyReLU with a negative slope of 0.2. During component separation, the input features are first mapped into a latent space using a shared encoder. Three dedicated decoders are then used to decode the three components. To achieve causal reasoning and hierarchical decoupling, two key mechanisms are introduced: first, an attention gating mechanism controls the flow of information to maintain appropriate independence between components; second, prior knowledge is introduced to guide the decoupling process. The reflective component is mainly related to the surface smoothness and the direction of the incident light. Therefore, during the decoding process, the spatial attention mechanism is used to highlight the highlight area. The oxidation component is related to the temperature distribution and oxidation time. During the decoding, the channel attention mechanism is used to enhance the features related to the color change. The forging component is related to the material deformation and internal structure. During the decoding, the multi-scale fusion mechanism is used to retain the detailed texture information.
[0048] Through this separation mechanism, the features at each scale level are component-decoupled, resulting in 15 decoupled feature maps (5 scale levels × 3 components). The dimensions of each decoupled feature map are the corresponding scale image size × 64 (feature dimension).
[0049] The feature discriminability is calculated for the decoupled features at each scale level, and the fusion weight coefficient is determined. Feature discriminability is assessed by calculating two metrics: feature variance and feature entropy. Feature variance reflects the dispersion of the feature distribution, while feature entropy reflects the uncertainty of the feature distribution. For each decoupled feature, the variance of its 64-dimensional feature vector is calculated to obtain the variance value. The feature histogram is then calculated (the feature values are divided into 10 uniform intervals), and the entropy of the histogram is then calculated. The variance and entropy values are normalized and weighted summed in a ratio of 0.5:0.5 to obtain the feature discriminability score.
[0050] Fusion weight coefficients are calculated based on feature discriminability. A softmax function is used to convert the discriminability scores into weights, so that the sum of the weights at all scale levels is 1. In practice, the weights of higher scale levels (such as the original resolution and the secondary resolution) are typically between 0.25 and 0.35, while the weights of lower scale levels are between 0.1 and 0.2, reflecting the importance of high-resolution features for contour extraction.
[0051] The decoupled features are weighted and combined using the fusion weight coefficients. The feature maps at each scale level are upsampled to the original resolution. The weighted summation is then performed according to the weight coefficients to obtain the fused features of the three components. In the actual implementation, the upsampling method uses bilinear interpolation to ensure a smooth transition of the feature maps.
[0052] The aluminum alloy tube target image was reconstructed based on the fused features. The reconstruction process used a convolutional neural network consisting of three convolutional layers with a 3×3 kernel size and 64, 32, and 3 channels (corresponding to the three RGB channels). The activation function was Reinforced Luminance (ReLU). The final layer used a sigmoid function to normalize the output values to the range of 0–1. The difference between the reconstructed image and the original image was calculated using the mean squared error (MSE) as the reconstruction loss.
[0053] At the same time, the spatial gradient of the decoupled features is calculated as a regularization constraint. The spatial gradient is calculated using the Sobel operator to calculate the horizontal and vertical gradients of the feature map. The L1 norm of the gradient magnitude is then calculated as the gradient loss. During training, the total loss function is a weighted sum of the reconstruction loss and the gradient loss, with a weight ratio of 1:0.2.
[0054] The parameters of the component separation unit were optimized based on the aforementioned loss function. The Adam optimizer was used, with an initial learning rate of 0.001, which was then reduced to 80% every 50 epochs. The batch size was set to 16, and training was performed for 200 epochs. The training data consisted of 500 images of an enhanced aluminum alloy tube target, 400 of which were used for training and 100 for validation. During training, validation performance was evaluated every 10 epochs, and the best-performing model parameters were saved.
[0055] After optimization is complete, the component separation unit outputs the contour feature data of the aluminum alloy tube target. Specifically, the forging component feature map is extracted, and threshold segmentation (the threshold value is set to 1.5 times the feature mean) is applied to identify high-response areas. Morphological operations (erosion followed by dilation, with a 3×3 rectangle as the structuring element) are then applied to remove noise. This results in the contour feature data of the aluminum alloy tube target. This contour feature data, including edge position coordinates and edge strength values, can be used for subsequent deformation analysis and quality assessment.
[0056] This method effectively separates the reflective, oxidized, and forged features of the aluminum alloy tube target surface through multi-scale decomposition and texture feature decoupling, overcoming the limitations of traditional methods in dealing with complex surface properties. This is particularly true in high-temperature forging environments, where aluminum alloy surfaces simultaneously exhibit multiple phenomena, including light reflection, oxidative discoloration, and deformation textures. This method accurately extracts forging-induced surface features while filtering out interfering factors such as reflective and oxidized features. This significantly improves the accuracy and stability of contour feature extraction, providing a reliable data foundation for subsequent quality assessment and process parameter optimization.
[0057] In an optional embodiment: The component separation unit realizes the decoupling and separation of the reflective component, the oxidation component, and the forging component through causal reasoning and hierarchical decoupling. The steps of obtaining the decoupling characteristics of each scale level include: The component separation unit includes a causal reasoning module and a hierarchical decoupling module; in the causal reasoning module, a causal graph structure is constructed between the reflective component, the oxidation component, and the forging component, and the conditional probability distribution between the components is calculated based on the causal graph structure to establish the mutual dependence relationship between the component features; in the hierarchical decoupling module, a decoupling loss function including a reliability assessment item is constructed, and the reliability assessment item calculates the decoupling confidence based on the feature distribution of the local area, and adopts an adaptive decoupling strength for different confidence areas; Based on the interdependence and the decoupling confidence, a progressive training strategy is adopted to perform feature decoupling, first establishing a decoupling benchmark in a high-confidence region, and then gradually extending the decoupling process to a low-confidence region to achieve reliable separation of the reflective component, the oxidized component, and the forged component; The feature reconstruction loss is calculated according to the decoupling benchmark, and the parameters of the causal reasoning module are optimized based on the feature reconstruction loss to obtain the decoupling features of each scale level.
[0058] Exemplarily, the component separation unit first receives multi-scale texture descriptors as input. For aluminum alloy tube targets, the texture descriptors at each scale level have a dimension of 64, corresponding to a native resolution of 1024×768×64, a second-level descriptor of 512×384×64, and so on. These descriptors contain rich information about the surface texture, but reflective, oxidized, and forged features are intermingled and require further separation.
[0059] In the causal inference module, a causal graph is constructed between the reflective, oxidized, and forged components. Based on physical prior knowledge, there is a clear causal relationship between the three components: the reflective component is primarily determined by surface geometry and incident light direction; the oxidized component is affected by temperature distribution and oxidation time; and the forged component is directly related to material flow and deformation. Therefore, in the constructed causal graph, the reflective component → oxidized component → forged component forms the primary causal link, while also considering the direct influence of the reflective component → forged component. The causal graph structure is represented by a directed graph network, with nodes corresponding to the three components and edges representing causal relationships. Each node contains a feature extractor consisting of three convolutional layers with a 3×3 kernel size and 64, 128, and 64 channels, respectively. Each layer is followed by batch normalization and LeakyReLU activation (negative slope of 0.2). The causal strength of edges is achieved through an attention gating mechanism. For example, for the edge from the reflective component to the oxidized component, an attention map (with the same dimensions as the feature map) is generated to control the strength of information transfer.
[0060] The conditional probability distribution between components is calculated based on the causal graph structure. For each component, its feature representation conditionally depends on its parent node component. For example, the oxidation component feature conditionally depends on the reflective component feature; the forging component feature conditionally depends on the reflective and oxidation components. The conditional probability distribution is estimated using variational inference methods, modeling each component's feature representation as a Gaussian distribution, with the distribution parameters (mean and variance) output by the feature extractor. In practice, for the oxidation component, its conditional distribution parameters are derived from the reflective component features via a fully connected layer; for the forging component, its conditional distribution parameters are derived from the concatenation of the reflective and oxidation component features via a fully connected layer. The interdependencies between component features are established through feature reconstruction and mutual information maximization. Feature reconstruction requires that the features of the child node components can be predicted from the features of the parent node component. For example, to reconstruct the oxidation component features from the reflective component features, the reconstruction network is a two-layer fully connected network with 128 hidden layer dimensions. Maximizing mutual information requires that related components share information and that unrelated components minimize mutual information. The specific implementation adopts the contrastive learning method, with the positive sample pairs as the feature pairs of the relevant components and the negative sample pairs as the feature pairs of the irrelevant components. The optimization is carried out through the InfoNCE loss function, and the temperature parameter is set to 0.07.
[0061] In the hierarchical decoupling module, a decoupling loss function is constructed that includes a reliability assessment term. Reliability assessment calculates the decoupling confidence based on the feature distribution of the local area. For each location in the image, the consistency of the features in the surrounding 9×9 neighborhood is calculated. The consistency is assessed by the cosine value of the angle between the feature vectors. The closer the cosine value is to 1, the more consistent the features are and the higher the decoupling confidence. In practical applications, the decoupling confidence is usually between 0.8-0.95 in smooth areas such as the cylindrical surface of the target tube body; the decoupling confidence is usually between 0.4-0.7 in areas with complex edges and textures such as forging wrinkles.
[0062] Adaptive decoupling strength is employed for regions with varying confidence levels. This strength is controlled by a regularization coefficient. For high-confidence regions (e.g., confidence > 0.8), a higher decoupling strength is employed, with a regularization coefficient of 1.0; for medium-confidence regions (e.g., confidence between 0.5 and 0.8), a regularization coefficient of 0.6; and for low-confidence regions (e.g., confidence < 0.5), a regularization coefficient of 0.3. This adaptive strategy strengthens decoupling in regions with clear feature distribution while maintaining appropriate flexibility in regions with ambiguous feature distribution.
[0063] Based on interdependencies and decoupling confidence, a progressive training strategy is employed for feature decoupling. Progressive training is divided into three phases: the first phase (the first 50 epochs) trains exclusively in high-confidence regions (confidence > 0.8) to establish a decoupling baseline; the second phase (51-100 epochs) expands to medium-confidence regions (confidence > 0.5); and the third phase (101-200 epochs) expands to the entire region. This strategy first establishes reliable decoupling patterns in regions with clear feature distributions and then transfers this decoupling knowledge to more complex regions. The network weights are updated differently in each training phase. In the first phase, only parameters associated with high-confidence regions are updated, freezing all other parameters; in the second phase, parameters associated with high- and medium-confidence regions are updated; and in the third phase, all parameters are updated. The learning rate is also adjusted in stages: 0.001 in the first phase, 0.0005 in the second phase, and 0.0001 in the third phase. The optimizer used is Adam, with momentum parameters beta1 and beta2 set to 0.9 and 0.999, respectively.
[0064] After reliably separating the reflective, oxidized, and forged components, the characteristic maps of each component clearly display surface characteristics of different physical origins. The reflective component feature map highlights areas of high gloss, such as smooth areas on the target surface; the oxidized component feature map highlights areas of color change, such as the oxide film formed in high-temperature areas; and the forged component feature map highlights areas of texture and deformation, such as streamlines and wrinkles formed during the forging process.
[0065] The feature reconstruction loss is calculated based on the decoupled benchmark. Feature reconstruction uses an encoder-decoder structure. The encoder maps the mixed features into three latent spaces, and the decoder reconstructs the three component features from each of the three latent spaces. The reconstructed features are then compared with the decoupled benchmark features, and the mean squared error is calculated as the reconstruction loss. The reconstruction loss weight is set to 1.0 for high-confidence regions, 0.7 for medium-confidence regions, and 0.4 for low-confidence regions.
[0066] In addition to the reconstruction loss, adversarial loss and mutual information loss are also introduced. The adversarial loss is implemented through a discriminator network, which consists of a four-layer convolutional network. It determines the authenticity of reconstructed features and decoupled baseline features, ensuring that the generated feature distribution is closer to the true distribution. Mutual information loss is implemented by minimizing the mutual information between uncorrelated components. For example, the mutual information between the reflective component and the forged component should be minimized, while the mutual information between correlated components should be within a certain range.
[0067] The parameters of the causal inference module were optimized based on the aforementioned loss function combination. The total loss function was a weighted sum of reconstruction loss, adversarial loss, and mutual information loss, with a weight ratio of 1:0.5:0.3. Training was performed using mini-batch gradient descent with a batch size of 16, and each epoch consisted of 25 batches. During training, model performance was evaluated on the validation set every 10 epochs. Training was terminated early if performance showed no improvement after five consecutive evaluations.
[0068] The optimized causal inference module accurately separates the reflective, oxidized, and forged components from the multi-scale texture descriptors, generating decoupled features at each scale level. For the original-resolution image of an aluminum alloy tube target, the decoupled features have dimensions of 1024 × 768 × 64 × 3 (width × height × number of channels × number of components); for the second-level image, the dimensions are 512 × 384 × 64 × 3, and so on. These decoupled features preserve the spatial structure and semantic information of the original features while separating their physical origins, providing a foundation for subsequent contour feature extraction and quality assessment.
[0069] This method effectively addresses the problem of mixed surface features in aluminum alloy tube targets through causal reasoning and hierarchical decoupling, accurately separating features of different physical origins. Compared to traditional methods, this approach fully considers the causal relationships and confidence distribution between features. It employs a progressive training strategy, extending the decoupling process from high-confidence regions to low-confidence regions, significantly improving the accuracy and robustness of the decoupling. This approach is particularly effective in extracting pure forging components from complex scenarios involving the simultaneous presence of surface reflection, oxidation, and deformation during the forging process.
[0070] In an optional embodiment: The steps of predicting deformation trends through graph structured calculation based on the topological relationship graph of the grid units and the contour feature data and generating forging parameter adjustment instructions include: Mapping the contour feature data of the aluminum alloy tube target to the topological relationship graph of the grid unit to construct a graph feature matrix, wherein the graph feature matrix includes geometric features and material state parameters of each grid node; A graph neural network is used to propagate features of the graph feature matrix, a message passing mechanism is established based on the connection relationship between grid nodes, and the stress distribution and deformation prediction value of each node are calculated through message aggregation. A deformation state vector is constructed based on the stress distribution and deformation prediction value. The deformation state vector includes the node displacement field, stress field and material flow characteristics, and is used to characterize the overall deformation state of the aluminum alloy tube target. Comparing the deformation state vector with a preset target shape, calculating the shape deviation, and predicting subsequent deformation trends based on the shape deviation and material deformation law; According to the predicted deformation trend, a forging parameter mapping function is established, wherein the forging parameter mapping function converts the predicted deviation into adjustment amounts of forging force, forging speed and temperature control parameters, and generates a forging parameter adjustment instruction.
[0071] For example, mapping the contour feature data of the aluminum alloy tube target to the topological relationship graph of the grid cells and constructing the graph feature matrix is a coherent process. First, the aluminum alloy tube target is discretized using a hexahedral grid. The standard tube target model is divided into 12 cells in the radial direction, 24 cells in the circumferential direction, and 36 cells in the axial direction, forming a total of 10,368 grid cells. These grid cells form the basic framework of the topological relationship graph. Each grid cell corresponds to a node in the graph, and the shared faces between adjacent cells correspond to edges in the graph. The grid topological relationship is represented by an adjacency matrix with a matrix dimension of 10,368 × 10,368. The element value 1 indicates adjacent and 0 indicates non-adjacent.
[0072] Contour feature data provides key observational information about surface nodes. A feature matching algorithm accurately locates these contour features to the corresponding nodes in the topological map. The matching process is based on the nearest neighbor principle in spatial coordinates. For each contour feature point, the Euclidean distance to the mesh surface node is calculated, and the feature is assigned to the node with the smallest distance. When a mesh node corresponds to multiple feature points, the feature values are fused using a weighted average method, with the weight inversely proportional to the distance.
[0073] For each surface node, the mapped features include the three-dimensional coordinates (X, Y, Z, in millimeters), the three components of the normal vector (normalized to the range [-1, 1]), the principal curvature value (in units of 1 / mm), and the forging component strength value (normalized to the range [0, 1]). For example, the feature data for a surface node in the middle of the target tube might be: coordinates (125.36, 0.00, 87.42) mm, normal vector (1.00, 0.00, 0.00), principal curvature of 0.008 / mm, and forging component strength of 0.65. To ensure mapping accuracy, bilinear interpolation is used to handle cases where contour feature points do not fully correspond to grid nodes, with the interpolation radius set to 1.5 times the grid cell size.
[0074] After the surface feature mapping is completed, the internal node features cannot be observed directly and need to be filled by propagating features from the surface to the interior. The propagation adopts a physics-based interpolation method to establish a feature transfer channel from the surface nodes to the internal nodes. Specifically, based on the basic principles of elastic-plastic mechanics and the surface deformation data and boundary conditions, Laplace interpolation is used to calculate the initial state estimate of the internal nodes. The Laplace interpolation solution is as follows: for each internal node, its eigenvalue is equal to the weighted average of the eigenvalues of all adjacent nodes, and the weight is inversely proportional to the distance. This ensures a smooth transition of the physical field, high computational efficiency and results that conform to physical meaning. For different types of physical fields, special propagation strategies are adopted: the temperature field is interpolated using the steady-state solution of the heat conduction equation, and the temperature distribution satisfies the Laplace equation; the stress field adopts an interpolation method constrained by the elastic equilibrium equation to ensure that the internal stress field meets the equilibrium condition; the deformation field is constrained by the displacement continuity condition to ensure the continuity and smoothness of the deformation field.
[0075] Through contour feature mapping and internal feature propagation, complete feature information is assigned to each node in the topological graph, thereby constructing a graph feature matrix. This matrix includes the geometric features and material state parameters of each mesh node. Geometric features include node coordinates (3D), displacement values (3D), and local deformation gradients (9D). Material state parameters include equivalent stress (1D), equivalent strain (1D), principal stress components (3D), temperature (1D), material flow velocity (3D) and flow direction (3D), hardening parameters (3D), and damage indicators (2D). The final graph feature matrix has dimensions of node number × feature dimension, which for the standard model is approximately 11200 × 32 (including surface and internal nodes). It fully records the geometric morphology and material state distribution of the aluminum alloy tube target in its current forging state.
[0076] A graph neural network is used to propagate features across the graph feature matrix. This approach employs a message-passing neural network (MPNN) architecture, encompassing four key steps: feature transformation, message generation, message aggregation, and state update. Feature transformation is achieved through a multilayer perceptron (MLP), which maps raw node features into a 128-dimensional hidden feature space. The MLP consists of two fully connected layers with ReLU activation function and a layer structure of 32-64-128 (input dimension - hidden dimension - output dimension).
[0077] The message generation process considers the interaction between node characteristics and those of adjacent nodes. For each pair of connected nodes, a feature difference vector is calculated and combined with edge attributes (such as inter-node distance and connection direction) to generate a message vector. The message generation function uses a three-layer fully connected network with a layer structure of 256-128-128 (input dimension after concatenation - hidden layer dimension - output dimension). For example, for adjacent nodes on a surface, the generated message contains information such as relative displacement, stress gradient, and temperature gradient to describe the local deformation state.
[0078] Message aggregation is achieved through a weighted summation mechanism. Weights are determined based on the distance between nodes and the direction of material flow. Nodes with closer distances and more consistent flow directions receive greater message weights. Weight calculation typically uses an attention mechanism. The attention score is calculated using a dot product followed by softmax normalization, with a temperature parameter set to 0.1. The aggregated message vector maintains a dimension of 128, incorporating information from all neighboring nodes.
[0079] State updates combine the current node state with aggregated messages to generate a new node state through a gated update mechanism. Similar to GRU units, gated updates contain reset and update gates that control the ratio of information retained and updated. The gating parameters are calculated based on node characteristics; typical values are: for nodes in high-stress areas, the update ratio is approximately 0.7-0.9; for nodes in low-stress areas, the update ratio is approximately 0.3-0.5. By repeatedly performing message passing and state updates (typically 3-5 iterations), graph neural networks are able to capture global deformation patterns.
[0080] After feature propagation is complete, the stress distribution and predicted deformation values for each node are calculated based on the connectivity between mesh nodes. The stress distribution is obtained by decoding node features and includes the three principal stress components and equivalent stress values. The predicted deformation value, including the node's displacement increment at future time steps, is predicted by the decoding network. The decoding network is a three-layer fully connected network with a layer structure of 128-64-16 (feature dimension - hidden layer dimension - output dimension). The activation function is Tanh, and the output is a 16-dimensional vector containing information such as the principal stress components, equivalent stress, displacement increment, and flow velocity.
[0081] A deformation state vector is constructed based on the stress distribution and deformation predictions. The deformation state vector is a compact representation of the overall deformation state of the aluminum alloy tube target. Key components are extracted from node features using principal component analysis (PCA). In implementation, the characteristic data of all nodes is first standardized, and then PCA is applied for dimensionality reduction, retaining the principal components required to explain 95% of the variance (typically 20-30 components). The deformation state vector contains three components: the node displacement field, the stress field, and the material flow characteristics. The displacement field describes the displacement direction and amplitude of each node; the stress field describes the stress state at each node; and the material flow characteristics describe the flow velocity and direction of the material. For example, the partial deformation state vector values for a typical forging state are: maximum displacement of 0.85 mm (occurring at the end of the tube target), maximum equivalent stress of 278 MPa (occurring at the center of the deformation zone), and maximum material flow velocity of 3.2 mm / s (in the radial direction).
[0082] The deformed state vector is compared to the preset target shape to calculate the shape deviation. The target shape is defined by a 3D CAD model, including the precise geometric dimensions of the target tube. Shape deviation is calculated using the minimum vertex-to-surface distance method. For each surface node, the distance from the target surface is calculated. Typical shape deviation metrics include maximum deviation (typically required to be within 0.5 mm), average deviation (typically required to be within 0.2 mm), and standard deviation (typically required to be within 0.15 mm).
[0083] Future deformation trends are predicted based on shape deviations and material deformation laws. Deformation trend prediction utilizes a combination of extrapolation and physical constraints. First, linear extrapolation is performed based on the current deformation rate to predict the position at future time points. Then, volume conservation and material constitutive relations are applied to correct the extrapolated results. This correction process accounts for material strengthening and temperature softening effects, and an iterative solution process ensures that the predicted results conform to physical laws. Deformation trends are quantitatively assessed using the deviation change rate and critical region identification methods. The deviation change rate is calculated based on the deviation change between two consecutive time steps, with positive values indicating increasing deviation and negative values indicating decreasing deviation. Critical regions are areas of shape deviation or stress concentration, identified by setting thresholds (e.g., deviation exceeding twice the average value or stress exceeding 90% of the material yield strength). Based on the predicted deformation trends, a forging parameter mapping function is established. This forging parameter mapping utilizes a combination of rule-based and table lookup methods to map shape deviations and deformation trends to specific forging parameter adjustments. The rule base, built based on forging expert experience and experimental data, contains processing strategies for different types of shape deviations. For example: When the radial expansion of the target tube end is too large (deviation > 0.3 mm), reduce the radial forging force by 5-10 kN and reduce the temperature of the area by 10-15 °C; When the radial shrinkage in the middle of the tube target is too large (deviation > 0.25 mm), increase the axial forging speed by 1-2 mm / s and increase the temperature in this area by 5-10°C. When uneven deformation occurs on the target surface (standard deviation > 0.2 mm), adjust the forging force distribution and increase the forging force in the area with insufficient deformation by 3-8 kN; When it is predicted that local stress concentration may occur (stress>300MPa), the current forging stage time is extended by 1-2 seconds and the forging speed is reduced by 2-3 mm / s.
[0084] The table lookup method enables fast parameter querying using a pre-calculated parameter mapping table. The mapping table dimensions are deviation type x deviation degree x current forging stage, and each table cell contains corresponding parameter adjustment suggestions. For example, for the "end radial expansion" type, a deviation of 0.35 mm, and the second forging stage, the table lookup suggests reducing the radial forging force by 8 kN, lowering the end temperature by 12°C, and reducing the forging speed by 1.5 mm / s.
[0085] The generated forging parameter adjustment instructions include adjustments for the upper die forging force (within a range of ±50kN), the lower die forging force (within a range of ±50kN), the radial forging speed (within a range of ±5mm / s), the axial forging speed (within a range of ±3mm / s), the temperature control parameters for the four zones (within a range of ±15°C for each zone), and the duration of the three forging stages (within a range of ±2s for each stage). These parameter adjustment instructions are transmitted to the forging equipment control system in a structured data format. Upon receiving the instructions, the equipment implements them step by step, and the results are verified through sensor feedback after each adjustment.
[0086] This method combines fundamental material forming theory with practical engineering experience to accurately predict deformation trends and optimize parameters in real time during the forging process of aluminum alloy tubular targets. For complex aluminum alloy tubular targets, this method, through a structured representation of material deformation and rule-based parameter mapping, enables foreseeing potential deformation defects and proactively adjusting process parameters.
[0087] In a second aspect, a forging system for an aluminum alloy tube target based on image processing is provided, comprising: The first unit is used to establish a three-dimensional digital model of the aluminum alloy tube target, divide the three-dimensional digital model into grid units, and establish a topological relationship diagram of the grid units; The second unit is used to use a dual-spectrum camera array to collect forging images of the aluminum alloy tube target from multiple perspectives, perform adaptive exposure compensation to obtain a multi-perspective fusion image; extract camera pose parameters from the multi-perspective fusion image to obtain corresponding perspective direction information; perform sinusoidal position encoding on the three-dimensional spatial coordinates of the grid unit and the perspective direction information, respectively, and input them into a neural radiation field network for volume rendering reconstruction to obtain a temperature radiation intensity mapping image; The third unit is configured to perform phase consistency registration and wavelet transform denoising on the temperature radiation intensity mapping image to obtain an enhanced aluminum alloy tube target image; input the enhanced aluminum alloy tube target image into a multi-scale texture decoupling network, obtain a local texture descriptor through multi-scale decomposition and texture feature extraction, input the local texture descriptor into a component separation unit, perform decoupling and separation of the reflection component, the oxidation component, and the forging component, and output the contour feature data of the aluminum alloy tube target; The fourth unit is used to predict the deformation trend through graph structured calculation based on the topological relationship diagram of the grid unit and the contour feature data, and generate a forging parameter adjustment instruction; and control the forging equipment according to the forging parameter adjustment instruction.
[0088] According to a third aspect, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
Claims
1. A forging method for an aluminum alloy tube target based on image processing, characterized in that: include: Establishing a three-dimensional digital model of the aluminum alloy tube target, dividing the three-dimensional digital model into grid units, and establishing a topological relationship diagram of the grid units; A dual-spectral camera array is used to capture forged images of an aluminum alloy tube target from multiple perspectives, and adaptive exposure compensation is performed to obtain a multi-perspective fusion image. Camera pose parameters are extracted from the multi-perspective fusion image to obtain corresponding viewing direction information. Sine position encoding is performed on the three-dimensional spatial coordinates of the grid cells and the viewing direction information, respectively, and the results are input into a neural radiation field network for volume rendering reconstruction to obtain a temperature radiation intensity mapping image. Phase consistency registration and wavelet transform denoising are performed on the temperature radiation intensity mapping image to obtain an enhanced aluminum alloy tube target image; the enhanced aluminum alloy tube target image is input into a multi-scale texture decoupling network, a local texture descriptor is obtained through multi-scale decomposition and texture feature extraction, the local texture descriptor is input into a component separation unit, a reflective component, an oxidation component, and a forging component are decoupled and separated, and contour feature data of the aluminum alloy tube target is output; Based on the topological relationship diagram of the grid unit and the contour feature data, the deformation trend is predicted through graph structured calculation to generate a forging parameter adjustment instruction; and the forging equipment is controlled according to the forging parameter adjustment instruction.
2. The method according to claim 1, characterized in that The steps of establishing a three-dimensional digital model of an aluminum alloy tube target, dividing the three-dimensional digital model into grid units, and establishing a topological relationship diagram of the grid units include: Establishing a three-dimensional NURBS surface model of an aluminum alloy tube target, wherein the three-dimensional NURBS surface model includes a cylindrical surface and an end surface, and the cylindrical surface and the end surface are constructed by control point coordinates, weight factors, and B-spline basis functions; Meshing the three-dimensional NURBS surface model, establishing an adaptive mesh density control function according to the curvature function, wherein the adaptive mesh density control function calculates a mesh unit size based on a minimum mesh size, a maximum mesh size, a directional weight coefficient, and a mesh density control parameter to generate tetrahedral mesh units; A topological relationship graph is constructed based on the tetrahedral grid unit, wherein the topological relationship graph includes a grid node set and a node connection edge set, wherein the grid node set is used as a vertex and the node connection edge set is used as an edge.
3. The method according to claim 1, characterized in that A dual-spectral camera array is used to capture forged images of an aluminum alloy tube target from multiple viewing angles, and adaptive exposure compensation is performed on the forged images to obtain a multi-view fused image. Camera pose parameters are extracted from the multi-view fused image to obtain corresponding viewing direction information. Sine position encoding is performed on the three-dimensional spatial coordinates of the grid cells and the viewing direction information, and the encoded features are input into a neural radiation field network. The steps of numerically integrating the neural radiation field network by a volume rendering method to obtain a temperature radiation intensity mapping image include: A dual-spectral camera array is used to collect visible light images and near-infrared images of the aluminum alloy tube target, and an adaptive weight function is established based on the scene brightness value to fuse the visible light image and the near-infrared image to obtain a multi-view fused image; Obtaining an intrinsic parameter matrix and an extrinsic parameter matrix of the dual-spectral camera array, calculating a camera optical center position based on the intrinsic parameter matrix and the extrinsic parameter matrix, constructing a viewing angle direction vector from the camera optical center position and an observation point position, and establishing a corresponding relationship between the multi-view fusion image and the viewing angle direction vector; Performing multi-layer sinusoidal position encoding on the three-dimensional spatial coordinates of the aluminum alloy tube target and the viewing angle direction vector, respectively, and inputting the encoded spatial coordinate features and viewing angle features into a neural radiation field network; The neural radiation field network is trained based on the multi-view fusion image and the corresponding view direction vector; the neural radiation field network outputs a volume density parameter and a direction-related radiation intensity parameter, numerically integrates the volume density parameter and the radiation intensity parameter along the direction of the light, calculates the cumulative transmittance, and obtains a continuous radiation field expression; The light cumulative radiation value is calculated based on the continuous radiation field expression, and the light cumulative radiation value is projected and mapped according to the viewing angle direction to obtain a temperature radiation intensity mapping image.
4. The method according to claim 3, characterized in that The step of training the neural radiation field network based on the multi-view fusion image and the corresponding view direction vector includes: Extracting and fusing features of the multi-view fusion image and the view direction vector to obtain a view perception feature; Constructing an adaptive sampling density function in three-dimensional space, wherein the adaptive sampling density function determines the sampling point density according to the angle between the surface normal vector of the sampling point and the viewing direction vector, and performing layered sampling along the ray direction based on the sampling point density to obtain a sampling point sequence; Input the view-aware features into a feature encoder, construct a multi-layer feature pyramid structure, and obtain a multi-scale feature representation through feature transformation and upsampling fusion; extract view-related features from the multi-scale feature representation and establish a feature-guided attention map; Inputting the sampling point sequence and the attention map into the neural radiance field network, and adaptively weighting the sampling point features based on the attention map; A composite loss function based on material properties is constructed, wherein the composite loss function includes a reconstruction loss term and a physical constraint loss term. The physical constraint loss term includes a thermal radiation continuity constraint and a surface oxide layer reflection characteristic constraint. The thermal radiation continuity constraint imposes a physical consistency constraint on the predicted radiation intensity distribution based on Planck's blackbody radiation law. The surface oxide layer reflection characteristic constraint constrains the reflection behavior under different viewing angles based on the optical properties of the high-temperature oxide film of aluminum alloy. The parameters of the neural radiation field network are optimized and trained based on the composite loss function.
5. The method according to claim 1, characterized in that The enhanced aluminum alloy tube target image is input into a multi-scale texture decoupling network, a local texture descriptor is obtained through multi-scale decomposition and texture feature extraction, the local texture descriptor is input into a component separation unit, and the reflective component, the oxidation component, and the forging component are decoupled and separated to output the contour feature data of the aluminum alloy tube target. The steps include: The enhanced aluminum alloy tube target image is down-sampled to generate images at multiple scale levels; local texture features are extracted from the images at the multiple scale levels, nonlinear feature transformation values are calculated within a local neighborhood of a preset size, and the nonlinear feature transformation values are weightedly combined with local weights to obtain a local texture descriptor for each scale level; Inputting the local texture descriptor into a component separation unit, the component separation unit realizes decoupling and separation of the reflective component, the oxidation component, and the forging component through causal reasoning and hierarchical decoupling, thereby obtaining decoupling features at each scale level; Calculating the feature discriminability of the decoupled features at each scale level, determining the corresponding fusion weight coefficient according to the feature discriminability, and performing weighted combination of the decoupled features using the fusion weight coefficient to obtain a fusion feature; An aluminum alloy tube target image is reconstructed based on the fusion feature, and a difference value between the reconstructed image and the original image and a spatial gradient value of the decoupling feature are calculated. The difference value and the spatial gradient value are used as constraints to optimize the parameters of the component separation unit, and the optimized aluminum alloy tube target contour feature data is output.
6. The method according to claim 5, characterized in that The component separation unit realizes the decoupling and separation of the reflective component, the oxidation component, and the forging component through causal reasoning and hierarchical decoupling. The steps of obtaining the decoupling characteristics of each scale level include: The component separation unit includes a causal reasoning module and a hierarchical decoupling module; in the causal reasoning module, a causal graph structure is constructed between the reflective component, the oxidation component, and the forging component, and the conditional probability distribution between the components is calculated based on the causal graph structure to establish the mutual dependence relationship between the component features; in the hierarchical decoupling module, a decoupling loss function including a reliability assessment item is constructed, and the reliability assessment item calculates the decoupling confidence based on the feature distribution of the local area, and adopts an adaptive decoupling strength for different confidence areas; Based on the interdependence and the decoupling confidence, a progressive training strategy is adopted to perform feature decoupling, a decoupling benchmark is established, and the decoupling process is gradually expanded to achieve reliable separation of the reflective component, the oxidation component, and the forging component; The feature reconstruction loss is calculated according to the decoupling benchmark, and the parameters of the causal reasoning module are optimized based on the feature reconstruction loss to obtain the decoupling features of each scale level.
7. The method according to claim 1, characterized in that The steps of predicting deformation trends through graph structured calculation based on the topological relationship graph of the grid units and the contour feature data and generating forging parameter adjustment instructions include: Mapping the contour feature data to a topological relationship graph of the grid cells to construct a graph feature matrix; A graph neural network is used to propagate features of the graph feature matrix. A message passing mechanism is established based on the connection relationship between grid nodes. The stress distribution and deformation prediction value of each node are calculated through message aggregation. A deformation state vector is constructed based on the stress distribution and deformation prediction values. The deformation state vector includes the node displacement field, stress field, and material flow characteristics. Comparing the deformation state vector with a preset target shape, calculating the shape deviation, and predicting subsequent deformation trends based on the shape deviation and material deformation law; According to the predicted deformation trend, a forging parameter mapping function is established, wherein the forging parameter mapping function converts the predicted deviation into adjustment amounts of forging force, forging speed and temperature control parameters, and generates a forging parameter adjustment instruction.
8. A forging system for an aluminum alloy tube target based on image processing, for implementing the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to establish a three-dimensional digital model of the aluminum alloy tube target, divide the three-dimensional digital model into grid units, and establish a topological relationship diagram of the grid units; The second unit is used to use a dual-spectrum camera array to collect forging images of the aluminum alloy tube target from multiple perspectives, perform adaptive exposure compensation to obtain a multi-perspective fusion image; extract camera pose parameters from the multi-perspective fusion image to obtain corresponding perspective direction information; perform sinusoidal position encoding on the three-dimensional spatial coordinates of the grid unit and the perspective direction information, respectively, and input them into a neural radiation field network for volume rendering reconstruction to obtain a temperature radiation intensity mapping image; The third unit is configured to perform phase consistency registration and wavelet transform denoising on the temperature radiation intensity mapping image to obtain an enhanced aluminum alloy tube target image; input the enhanced aluminum alloy tube target image into a multi-scale texture decoupling network, obtain a local texture descriptor through multi-scale decomposition and texture feature extraction, input the local texture descriptor into a component separation unit, perform decoupling and separation of the reflection component, the oxidation component, and the forging component, and output the contour feature data of the aluminum alloy tube target; The fourth unit is used to predict the deformation trend through graph structured calculation based on the topological relationship diagram of the grid unit and the contour feature data, and generate a forging parameter adjustment instruction; and control the forging equipment according to the forging parameter adjustment instruction.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Online quality monitoring method and system based on optical multispectral fusion
CN119198566A
Real-time detection system in forging forming process and three-dimensional reconstruction method thereof
CN119273840A
Titanium alloy forging process optimization method and system based on image processing
CN119810091A
Dynamic three-dimensional scene reconstruction and rendering system, method and related equipment
CN120198601A
Three-dimensional radiation field modeling method for performing infrared factor reasoning through poses
CN120431250A
Cited By
Gas leakage three-dimensional reconstruction method and system based on thermo-optic fusion
CN121170165A
Forging parameter self-adaptive adjusting method and system
CN121300078A
Industrial mechanism simulation three-dimensional image rendering method based on domestic environment
CN122223206A