A method and system for editing abnormal areas in railway freight train images
Through the abnormal area editing method of railway freight train images, using large language model and pre-trained model decomposition and fusion technology, the lag and high cost problems of railway freight safety inspection are solved, and efficient and accurate abnormal area detection and image quality improvement are achieved.
Patent Information
- Application Number
- CN202510969470.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-15
AI Technical Summary
The existing railway freight safety inspection technology has problems of lag and high cost, and mainly relies on traditional manual monitoring, which cannot meet the requirements of real-time and economy.
A method for editing abnormal areas in railway freight train images is adopted. The editing instructions are decomposed into subtask chains through a large language model. The structured semantic parsing network and pre-trained model are combined to extract target semantic information and image features, generate spatial masks and semantic control vectors, perform image editing and fusion optimization, and generate edited images of abnormal areas.
It achieves efficient and accurate detection of abnormal areas in railway freight train images, reduces costs, improves detection efficiency, avoids the lag and high cost of manual monitoring, and significantly improves the quality and credibility of the generated images.
Smart Images

Figure CN120472051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for editing abnormal areas in railway freight train images. Background Art
[0002] Railway transportation plays an important role in the national economy and people's lives. With the development of globalization and rapid economic growth, the volume of railway freight transportation continues to increase, and the demand for efficient and accurate cargo inspection is increasing. Ensuring railway transportation safety is of great significance to social stability and development.
[0003] At present, in the field of railway freight safety inspection, the main problems are the presence of foreign objects inside the carriage, whether the ropes of the tarpaulin outside the carriage are fastened, whether the doors and windows of the carriage are closed, etc. Traditional manual monitoring is usually adopted, which has serious lags and high labor costs. It can no longer meet the dual requirements of modern railway transportation for real-time and economy. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a method and system for editing abnormal areas in railway freight train images, aiming to solve the technical problems in the existing technology of traditional manual monitoring of railway freight safety, which has serious lags and high costs.
[0005] One aspect of the present invention is to provide a method for editing abnormal areas in railway freight train images, the method comprising:
[0006] Obtaining original images of a railway freight train carriage and editing instructions for abnormal areas;
[0007] Inputting the editing instruction and the original image into a large language model for parsing, and then gradually decomposing the editing instruction into a subtask chain through a thought chain reasoning mechanism, wherein the subtask chain includes editing object positioning, abnormal content construction, and regional fusion optimization;
[0008] Extracting target semantic information from the editing instruction, then extracting image features from the original image, locating a target area of the image features according to the target semantic information, generating a spatial mask, and outputting the target area mask to achieve editing object positioning;
[0009] fusing the target semantic information and the image features to obtain a semantic control vector, performing image editing based on the semantic control vector, the spatial mask, and the image features to obtain an initial edited image, thereby constructing abnormal content;
[0010] The target area mask is preprocessed to obtain a fusion boundary area, and the original image and the initial edited image are matched and fused in the fusion boundary area to obtain an abnormal area edited image, thereby achieving regional fusion optimization.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: through the railway freight train image abnormal area editing method provided by the present invention, the editing instructions are gradually decomposed into subtask chains, and the subtask chains include editing object positioning, abnormal content construction and regional fusion optimization, that is, decomposed into sub-steps that can be executed step by step, reducing the complexity of image editing, and by converting the editing instructions into triples and extracting target semantic information, while combining the image features of the original image to locate the target area, it is possible to more accurately understand the editing intention and accurately find the object to be edited. The generated spatial mask can clearly mark the editing target area, delineate clear boundaries for subsequent editing operations, complete the editing object positioning, and then based on the semantic control vector, spatial mask and image The image is edited based on the features to complete the construction of abnormal content. The generated initial edited image can make the abnormal content better adapt to the original image at the semantic and feature levels, thereby improving the quality and credibility of the content after image editing. Then, by matching and fusing the initial edited image, the transition traces of boundaries, textures, and lighting / colors are reduced, and the edited image of the abnormal area with optimized visual consistency is output to achieve regional fusion optimization, which can effectively improve the quality of the edited image of the abnormal area and the efficiency of image editing. Abnormal areas of railway freight train images can be detected without manual monitoring, reducing costs and improving efficiency, thereby solving the technical problems of traditional manual monitoring of railway freight safety in the existing technology, which has serious lags and high costs.
[0012] According to one aspect of the above technical solution, the steps of extracting target semantic information from the editing instruction, extracting image features from the original image, locating the target area of the image features according to the target semantic information, and generating a spatial mask specifically include:
[0013] The editing instruction is input into the structured semantic parsing network to obtain a triplet, which includes the editing target object, the editing operation type, and the editing space location, and is expressed as:
[0014] ,
[0015] in, is a triple, To edit the target object, For edit operation type, Locate the editing space;
[0016] The triplet is input into the pre-trained text editor BERT to extract the target semantic information. The calculation formula is:
[0017] ,
[0018] in, is the target semantic information, The calculation function for the pre-trained file editor BERT, for dimensional real vector space;
[0019] The original image is input into the pre-trained SAM editor to extract image features, which are expressed as:
[0020] ,
[0021] in, is the image feature, is the calculation function of the pre-trained SAM editor, is the original image, for dimensional real vector space, 、 、 are the height, width, and number of channels of the image features respectively;
[0022] The target semantic information and the image features are mapped to the same vector space, and the similarity between the image features and the target semantic information is calculated for each spatial position to obtain a heat map. The calculation formula is:
[0023] ,
[0024] in, For image features Feature information of linear projection on position, is the linear projection matrix of image features, For image features Feature information on the location, for dimensional real vector space, is the semantic feature of the target semantic information on the linear projection, is the linear projection matrix of the target semantic information, Heat map exist The similarity in position, , ;
[0025] The heat map is binarized to obtain a spatial mask, which is calculated as follows:
[0026] ,
[0027] in, is the spatial mask exist The binary content of the position, Represents a spatial mask The position distribution of the editing target area, 1 represents the editing target area, 0 represents the non-editing target area, is the heatmap threshold, .
[0028] According to one aspect of the above technical solution, the step of constructing the target area mask specifically includes:
[0029] The spatial mask is upsampled using a bilinear interpolation algorithm, and the size of the spatial mask is restored to the size of the original image to obtain a target area mask.
[0030] According to one aspect of the above technical solution, the step of fusing the target semantic information with the image features to obtain a semantic control vector specifically includes:
[0031] The target semantic information and the image features are fused through a cross-attention mechanism to obtain a semantic control vector, which is calculated as follows:
[0032] ,
[0033] in, 、 、 is a learnable matrix, is the activation function, is the query matrix, is the bond matrix, is the value matrix, is the semantic control vector, is the transpose of the key matrix, are learnable parameters, is the calculation function of the cross attention mechanism.
[0034] According to one aspect of the above technical solution, the step of performing image editing based on the semantic control vector, the spatial mask, and the image features to obtain an initial edited image and implement abnormal content construction specifically includes:
[0035] The original image is initialized with noise in the area of the spatial mask to obtain an initial noise image. The calculation formula is as follows:
[0036] ,
[0037] in, is the pixel point in the initial noise image content, To edit the target area, is the pixel point in the spatial mask content, is the pixel in the original image content, represents the standard normal noise initialization function;
[0038] The initial noisy image is used as the starting image of the diffusion process. Based on the image features, the semantic control vector and the spatial mask, diffusion denoising is performed step by step to obtain the initial edited image. The calculation formula is as follows:
[0039] ,
[0040] in, is the time step The predicted denoised image, is the calculation function of the noise prediction model, is the time step The predicted denoised image until =0, get the initial edited image.
[0041] According to one aspect of the above technical solution, the step of preprocessing the target area mask to obtain the fused boundary area specifically includes:
[0042] ,
[0043] in, To fuse the boundary area, is the target area mask, for The radius is The expansion operation, for The radius is corrosion operation.
[0044] According to one aspect of the above technical solution, the steps of matching and fusing the original image and the initial edited image in the fusion boundary region to obtain the abnormal region edited image and realizing regional fusion optimization specifically include:
[0045] Texture matching is performed on the original image and the initial edited image in the fusion boundary area, and edge smoothing is performed using a bilateral filter. The calculation formula is:
[0046] ,
[0047] in, Pixels in the initial edited image for edge smoothing content, is the normalization factor, Pixel Neighborhood, is the pixel in the initial edited image content, is the pixel in the original image content, Pixel Neighborhood pixels of is the spatial Gaussian filter calculation function, is the pixel Gaussian filter calculation function, pixel point Belongs to the fusion boundary area;
[0048] Based on the original image, the initial edited image is subjected to illumination correction, and the calculation formula is:
[0049] ,
[0050] in, The pixels of the initial edited image after lighting correction content, The pixels of the original edited image without light correction content, Calculate function for illumination correction;
[0051] Based on the original image, color correction is performed on the initial edited image after illumination correction, and the calculation formula is:
[0052]
[0053] in, represents the standard deviation on the feature channel, represents the mean value on the feature channel, Compute function for adaptive instance normalization;
[0054] The corrected initial edited image is fused with the original image to obtain an edited image of the abnormal area.
[0055] According to one aspect of the above technical solution, the corrected initial edited image is fused with the original image to obtain the abnormal area edited image using the following calculation formula:
[0056] ,
[0057] in, Edit the image for abnormal areas, is the original image, is the initial edited image after correction.
[0058] According to one aspect of the above technical solution, the method further includes:
[0059] The user interacts by inputting text descriptions, parsing and updating to generate new target semantic information, replacing the original image with the initial edited image, adjusting the target area positioning, and regenerating a new edited image of the abnormal area;
[0060] Multiple rounds of interaction are performed until a new edited image of the abnormal area that meets user needs is generated.
[0061] Another aspect of the present invention is to provide a railway freight train image abnormal area editing system, the system is used to implement the above-mentioned railway freight train image abnormal area editing method, the system comprising:
[0062] An editing instruction acquisition module is used to acquire an original image of a railway freight train carriage and editing instructions for abnormal areas;
[0063] A subtask decomposition module is used to input the editing instructions and the original image into a large language model for parsing, and then gradually decompose the editing instructions into subtask chains through a thought chain reasoning mechanism. The subtask chains include editing object positioning, abnormal content construction, and regional fusion optimization;
[0064] An editing object positioning module is used to extract target semantic information from the editing instruction, then extract image features from the original image, locate the target area of the image features according to the target semantic information, generate a spatial mask, and output the target area mask to achieve editing object positioning;
[0065] an abnormal content construction module, configured to fuse the target semantic information with the image features to obtain a semantic control vector, perform image editing based on the semantic control vector, the spatial mask, and the image features, obtain an initial edited image, and implement abnormal content construction;
[0066] The regional fusion optimization module is used to pre-process the target area mask to obtain a fusion boundary area, match and fuse the original image with the initial edited image in the fusion boundary area to obtain an abnormal area edited image, and realize regional fusion optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0068] Figure 1 This is a flow chart of a method for editing abnormal regions in a railway freight train image according to the first embodiment of the present invention;
[0069] Figure 2 This is a structural block diagram of a railway freight train image abnormal area editing system in a second embodiment of the present invention;
[0070] Component symbol description in the attached figure:
[0071] Editing instruction acquisition module 100, subtask decomposition module 200, editing object positioning module 300, abnormal content construction module 400, and regional fusion optimization module 500. DETAILED DESCRIPTION
[0072] To make the objectives, features, and advantages of the present invention more readily apparent, the following detailed description of specific embodiments of the present invention is provided in conjunction with the accompanying drawings. The accompanying drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.
[0073] Example 1
[0074] See also Figure 1 , shown is a method for editing abnormal regions in railway freight train images provided by a first embodiment of the present invention, the method comprising steps S10 to S14:
[0075] Step S10, obtaining an original image of a railway freight train carriage and an editing instruction for an abnormal area;
[0076] By way of example and not limitation, the editing instruction may be an editing instruction input by a user or an editing instruction defined by the system.
[0077] Furthermore, the original image may include any one of the three perspective images of the front, back, and top of the carriage of the railway freight train.
[0078] Step S11: Input the editing instruction and the original image into a large language model for parsing, and then gradually decompose the editing instruction into a subtask chain through a thought chain reasoning mechanism. The subtask chain includes editing object positioning, abnormal content construction, and regional fusion optimization;
[0079] Furthermore, the editing instructions and the original image are parsed by a large language model, such as the multimodal large model LLaVA.
[0080] For example, rather than limitation, editing instructions and original images are jointly modeled based on text prompts and image perception, and the output structured editing intention unit can be defined as a function:
[0081] ,
[0082] in, is the original image, To edit the command, For the identified target objects, For the semantic information or action requirement information of the target object, is the total number of target objects.
[0083] Furthermore, the editing instruction is gradually decomposed into subtask chains through the thought chain reasoning mechanism, which is expressed as:
[0084] ,
[0085] in, For the subtask chain, Indicates the Subtask, Each subtask depends on the output of the previous subtask, that is, , guided by the joint editing instructions and the original image, the dependency can be expressed as:
[0086] ,
[0087] in, For the An operation module for each subtask.
[0088] Furthermore, to improve robustness, when an editing instruction is ambiguous or lacks details, the system will automatically generate possible supplementary editing instructions based on historical editing instructions and confirm with the user.
[0089] The editing instructions are gradually decomposed into sub-task chains, that is, decomposed into sub-steps that can be executed step by step, and the sub-tasks are executed in sequence to achieve sequential dependency and collaboration, thereby reducing the complexity of image editing.
[0090] Step S12, extracting target semantic information from the editing instruction, then extracting image features from the original image, locating the target area of the image features according to the target semantic information, generating a spatial mask, and outputting the target area mask to achieve editing object positioning;
[0091] Specifically, the editing instruction is input into the structured semantic parsing network to obtain a triplet, which includes the editing target object, the editing operation type, and the editing space location, and is expressed as:
[0092] ,
[0093] in, is a triple, To edit the target object, For edit operation type, Locate the editing space;
[0094] The triplet is input into the pre-trained text editor BERT to extract the target semantic information. The calculation formula is:
[0095] ,
[0096] in, is the target semantic information, The calculation function for the pre-trained file editor BERT, for dimensional real vector space;
[0097] Furthermore, the text encoder BERT (Bidirectional Encoder Representations from Transformers) is a bidirectional encoder representation pre-trained language model based on the Transformer architecture.
[0098] The original image is input into the pre-trained SAM editor to extract image features, which are expressed as:
[0099] ,
[0100] in, is the image feature, is the calculation function of the pre-trained SAM editor, is the original image, for dimensional real vector space, 、 、 are the height, width, and number of channels of the image features respectively;
[0101] Furthermore, the pre-trained encoder in the SAM editor (Segment Anything Model) extracts image features including multi-scale semantic and edge texture information at each location in the original image, which facilitates the subsequent generation of spatial masks.
[0102] The target semantic information and the image features are mapped to the same vector space, and the similarity between the image features and the target semantic information is calculated for each spatial position to obtain a heat map. The calculation formula is:
[0103] ,
[0104] in, For image features Feature information of linear projection on position, is the linear projection matrix of image features, For image features Feature information on the location, for dimensional real vector space, is the semantic feature of the target semantic information on the linear projection, is the linear projection matrix of the target semantic information, Heat map exist The similarity in position, , ;
[0105] The heat map is binarized to obtain a spatial mask, which is calculated as follows:
[0106] ,
[0107] in, is the spatial mask exist The binary content of the position, Represents a spatial mask The position distribution of the editing target area, 1 represents the editing target area, 0 represents the non-editing target area, is the heatmap threshold, .
[0108] Furthermore, 0 represents a non-editing target area, and the image content remains unchanged.
[0109] In addition, the steps of constructing the target area mask include:
[0110] The spatial mask is upsampled using a bilinear interpolation algorithm, and the size of the spatial mask is restored to the size of the original image to obtain a target area mask.
[0111] It should be noted that by converting editing instructions into triplets and extracting target semantic information, and combining the image features of the original image to locate the target area, it is possible to more accurately understand the editing intention, accurately find the objects that need to be edited, avoid indiscriminate modifications, and improve the targeted nature of editing.
[0112] Furthermore, the generated spatial mask clearly marks the editing target area, defining clear boundaries for subsequent editing operations, ensuring that the editing content only acts on the editing target area, minimizing interference with other parts of the image, making the editing results meet expectations in terms of spatial scope, and enhancing the accuracy and reliability of editing.
[0113] Step S13: fusing the target semantic information and the image features to obtain a semantic control vector, performing image editing based on the semantic control vector, the spatial mask, and the image features to obtain an initial edited image, thereby constructing abnormal content.
[0114] In order to enhance semantic consistency, the target semantic information and the image features are fused through a cross-attention mechanism to obtain a semantic control vector, which is calculated as follows:
[0115] ,
[0116] in, 、 、 is a learnable matrix, is the activation function, is the query matrix, is the bond matrix, is the value matrix, is the semantic control vector, is the transpose of the key matrix, are learnable parameters, is the calculation function of the cross attention mechanism.
[0117] Further, Used to balance the magnitude of the dot product to prevent The gradient of the function vanishes.
[0118] It should be noted that the semantic control vector is generated through the cross-attention mechanism to achieve the association and fusion between features, so as to improve the ability to capture the contextual semantics of the original image and provide feature-level guidance for subsequent image editing.
[0119] Furthermore, the step of obtaining the initial edited image includes:
[0120] The original image is initialized with noise in the area of the spatial mask to obtain an initial noise image. The calculation formula is as follows:
[0121] ,
[0122] in, is the pixel point in the initial noise image content, To edit the target area, is the pixel point in the spatial mask content, is the pixel in the original image content, represents the standard normal noise initialization function;
[0123] The initial noisy image is used as the starting image of the diffusion process. Based on the image features, the semantic control vector and the spatial mask, diffusion denoising is performed step by step to obtain the initial edited image. The calculation formula is as follows:
[0124] ,
[0125] in, is the time step The predicted denoised image, is the calculation function of the noise prediction model, is the time step The predicted denoised image until =0, get the initial edited image.
[0126] It should be noted that the initial noise image is , at each diffusion time step , input the current predicted denoised image , image features , semantic control vector and spatial mask , predict the next diffusion time step The predicted denoised image.
[0127] It should be noted that the editing of the initial edited image is based on the diffusion generation architecture, that is, the semantic control vector and spatial mask are added in the noise reconstruction process based on the StableDiffusion model to realize the editing of the target area.
[0128] Furthermore, image editing is performed based on semantic control vectors, spatial masks, and image features. The resulting initial edited image can better align the anomalous content (understood as content that complies with the editing instructions) with the original image at the semantic and feature levels, making the anomalous content appear natural and reasonable, thereby improving the quality and credibility of the edited content. For example, the semantically generated new elements are consistent with the original image in terms of style and texture.
[0129] Secondly, the layered processing method of constructing abnormal content after locating the editing object makes the editing process clear and orderly. Positioning provides the basis for construction, and construction is expanded based on the positioning results. The combination of the two makes the construction of abnormal content both accurate and able to ensure the degree of integration with the original image. Compared with simple and crude editing methods, it can produce better quality and more demand-oriented editing results.
[0130] Step S14 , preprocessing the target area mask to obtain a fusion boundary area, matching and fusing the original image and the initial edited image in the fusion boundary area to obtain an abnormal area edited image, thereby achieving regional fusion optimization.
[0131] Specifically, the target area mask is morphologically expanded and eroded to obtain the fusion boundary area. The calculation formula is as follows:
[0132] ,
[0133] in, To fuse the boundary area, is the target area mask, for The radius is The expansion operation, for The radius is corrosion operation.
[0134] It should be noted that the fusion boundary area is a transition area where the boundary of the target area mask extends outward, so as to avoid the awkward feeling of directly splicing the edited target area with the original image.
[0135] At the same time, the edges of the fusion boundary area are smoothed by using a bilateral filter to reduce edge mutations and make the texture transition between the editing target area and the original image more natural. Specifically:
[0136] Texture matching is performed on the original image and the initial edited image in the fusion boundary area, and edge smoothing is performed using a bilateral filter. The calculation formula is:
[0137] ,
[0138] in, Pixels in the initial edited image for edge smoothing content, is the normalization factor, Pixel Neighborhood, is the pixel in the initial edited image content, is the pixel in the original image content, Pixel Neighborhood pixels of is the spatial Gaussian filter calculation function, is the pixel Gaussian filter calculation function, pixel point Belongs to the fusion boundary area.
[0139] Furthermore, in order to achieve consistency in lighting and color between the target editing area and the original image and unify the color style, it is necessary to perform lighting and color correction on the target editing area. Specifically:
[0140] Based on the original image, the initial edited image is subjected to illumination correction, and the calculation formula is:
[0141] ,
[0142] in, The pixels of the initial edited image after lighting correction content, The pixels of the original edited image without light correction content, Calculate function for illumination correction;
[0143] Based on the original image, color correction is performed on the initial edited image after illumination correction, and the calculation formula is:
[0144]
[0145] in, represents the standard deviation on the feature channel, represents the mean value on the feature channel, Compute function for adaptive instance normalization;
[0146] The corrected initial edited image is fused with the original image to obtain the edited image of the abnormal area. The calculation formula is:
[0147] ,
[0148] in, Edit the image for abnormal areas, is the original image, is the initial edited image after correction.
[0149] It should be noted that in order to achieve regional fusion optimization, the initial edited image is matched and fused to reduce the transition traces of boundaries, textures, and lighting / colors, and the edited image of the abnormal area with optimized visual consistency is output to improve image quality and enhance realism.
[0150] In addition, the method further comprises:
[0151] The user interacts by inputting text descriptions, parsing and updating to generate new target semantic information, replacing the original image with the initial edited image, adjusting the target area positioning, and regenerating a new edited image of the abnormal area;
[0152] Multiple rounds of interaction are performed until a new edited image of the abnormal area that meets user needs is generated.
[0153] It should be noted that interactive iterative editing allows users to precisely control the editing effects, achieve fine-grained image re-editing, improve editing efficiency, and optimize user experience.
[0154] Compared with the prior art, the railway freight train image abnormal area editing method shown in this embodiment reduces the complexity of image editing by gradually decomposing the editing instructions into a subtask chain, which includes editing object positioning, abnormal content construction and area fusion optimization, that is, decomposing into sub-steps that can be executed step by step. By converting the editing instructions into triples and extracting the target semantic information, and combining the image features of the original image to locate the target area, the editing intention can be understood more accurately and the object to be edited can be found accurately. The generated spatial mask can clearly mark the editing target area, delineate clear boundaries for subsequent editing operations, complete the editing object positioning, and then perform the editing operation based on the semantic control vector, spatial mask and image features. Image editing is performed to complete the construction of abnormal content. The generated initial edited image can make the abnormal content better adapt to the original image at the semantic and feature levels, thereby improving the quality and credibility of the content after image editing. Then, by matching and fusing the initial edited image, the transition traces of boundaries, textures, and lighting / colors are reduced, and the edited image of the abnormal area with optimized visual consistency is output to achieve regional fusion optimization, which can effectively improve the quality of the edited image of the abnormal area and the efficiency of image editing. Abnormal areas of railway freight train images can be detected without manual monitoring, reducing costs and improving efficiency, thereby solving the technical problems of traditional manual monitoring of railway freight safety in the existing technology, which has serious lags and requires high costs.
[0155] Example 2
[0156] See also Figure 2 , shown is a railway freight train image abnormal area editing system provided by a second embodiment of the present invention, the system comprising:
[0157] The editing instruction acquisition module 100 is used to acquire the original image of the railway freight train carriage and the editing instructions of the abnormal area;
[0158] A subtask decomposition module 200 is configured to input the editing instruction and the original image into a large language model for parsing, and then gradually decompose the editing instruction into a subtask chain through a thought chain reasoning mechanism. The subtask chain includes editing object positioning, abnormal content construction, and regional fusion optimization;
[0159] The editing object positioning module 300 is used to extract target semantic information from the editing instruction, extract image features from the original image, locate the target area of the image features according to the target semantic information, generate a spatial mask, and output the target area mask to achieve editing object positioning;
[0160] Abnormal content construction module 400, configured to fuse the target semantic information with the image features to obtain a semantic control vector, perform image editing based on the semantic control vector, the spatial mask, and the image features to obtain an initial edited image, thereby constructing abnormal content;
[0161] The regional fusion optimization module 500 is used to pre-process the target region mask to obtain a fusion boundary region, match and fuse the original image with the initial edited image in the fusion boundary region to obtain an abnormal region edited image, and realize regional fusion optimization.
[0162] Compared with the prior art, the railway freight train image abnormal area editing system shown in this embodiment decomposes the system into progressively executable sub-steps through a subtask decomposition module, reducing the complexity of image editing. The editing object positioning module converts editing instructions into triples and extracts target semantic information. Simultaneously, the target area is positioned based on the image features of the original image, enabling a more accurate understanding of the editing intent and precise identification of the object to be edited. The generated spatial mask clearly marks the editing target area, defining a clear boundary for subsequent editing operations. The abnormal content construction module performs image editing based on the semantic control vector, spatial mask, and image features. The generated initial edited image can better match the abnormal content with the original image at the semantic and feature levels, improving the quality and credibility of the edited image content. The initial edited image is then matched and fused through a regional fusion optimization module to reduce boundary, texture, and lighting / color transition traces. This effectively improves the quality of the edited image of the abnormal area and enhances the efficiency of image editing. Abnormal areas in railway freight train images can be detected without manual monitoring, reducing costs and improving efficiency. This solves the technical problems of traditional manual monitoring of railway freight safety in the prior art, which suffers from severe lags and high costs.
[0163] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0164] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as a sequenced list of executable instructions for implementing logical functions, and may be embodied in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can execute instructions).
[0165] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0166] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for editing abnormal areas in railway freight train images, characterized in that: The method comprises: Obtaining original images of a railway freight train carriage and editing instructions for abnormal areas; Inputting the editing instruction and the original image into a large language model for parsing, and then gradually decomposing the editing instruction into a subtask chain through a thought chain reasoning mechanism, wherein the subtask chain includes editing object positioning, abnormal content construction, and regional fusion optimization; Extracting target semantic information from the editing instruction, then extracting image features from the original image, locating a target area for the image features according to the target semantic information, generating a spatial mask, and outputting the target area mask to achieve editing object positioning, including: The editing instruction is input into the structured semantic parsing network to obtain a triplet, which includes the editing target object, the editing operation type, and the editing space location, and is expressed as: , in, is a triple, To edit the target object, For edit operation type, To locate the editing space, The triples are input into the pre-trained text editor BERT to extract the target semantic information. The calculation formula is: , in, is the target semantic information, The calculation function for the pre-trained file editor BERT, for dimensional real vector space, The original image is input into the pre-trained SAM editor to extract image features, which are expressed as: , in, is the image feature, is the calculation function of the pre-trained SAM editor, is the original image, for dimensional real vector space, 、 、 are the height, width, and number of channels of the image features, respectively. The target semantic information and the image features are mapped to the same vector space, and the similarity between the image features and the target semantic information is calculated for each spatial position to obtain a heat map. The calculation formula is: , in, For image features Feature information of linear projection on position, is the linear projection matrix of image features, For image features Feature information on the location, for dimensional real vector space, is the semantic feature of the target semantic information on the linear projection, is the linear projection matrix of the target semantic information, Heat map exist The similarity in position, , , The heat map is binarized to obtain a spatial mask, which is calculated as follows: , in, Space mask exist The binary content of the position, Represents a spatial mask The position distribution of the editing target area, 1 represents the editing target area, 0 represents the non-editing target area, is the heatmap threshold, ; The target semantic information and the image features are fused to obtain a semantic control vector, and image editing is performed based on the semantic control vector, the spatial mask, and the image features to obtain an initial edited image, thereby achieving abnormal content construction, including: The original image is initialized with noise in the area of the spatial mask to obtain an initial noise image. The calculation formula is as follows: , in, is the pixel point in the initial noise image content, To edit the target area, is the pixel point in the spatial mask content, is the pixel in the original image content, represents the standard normal noise initialization function, The initial noisy image is used as the starting image of the diffusion process. Based on the image features, the semantic control vector and the spatial mask, diffusion denoising is performed step by step to obtain the initial edited image. The calculation formula is as follows: , in, is the time step The predicted denoised image, is the calculation function of the noise prediction model, is the time step The predicted denoised image until =0, get the initial edited image; The target area mask is preprocessed to obtain a fusion boundary area, and the original image and the initial edited image are matched and fused in the fusion boundary area to obtain an abnormal area edited image, thereby achieving regional fusion optimization.
2. The method for editing abnormal areas of railway freight train images according to claim 1, characterized in that: The steps of constructing the target area mask include: The spatial mask is upsampled using a bilinear interpolation algorithm, and the size of the spatial mask is restored to the size of the original image to obtain a target area mask.
3. The method for editing abnormal areas of railway freight train images according to claim 1, characterized in that: The step of fusing the target semantic information and the image features to obtain a semantic control vector specifically includes: The target semantic information and the image features are fused through a cross-attention mechanism to obtain a semantic control vector, which is calculated as follows: , in, 、 、 is a learnable matrix, is the activation function, is the query matrix, is the bond matrix, is the value matrix, is the semantic control vector, is the transpose of the key matrix, are learnable parameters, is the calculation function of the cross attention mechanism.
4. The method for editing abnormal areas of railway freight train images according to claim 2, characterized in that: The step of preprocessing the target area mask to obtain a fused boundary area specifically includes: Perform morphological dilation and erosion on the target area mask to obtain the fused boundary area. The calculation formula is as follows: , in, To fuse the boundary area, is the target area mask, for The radius is The expansion operation, for The radius is corrosion operation.
5. The method for editing abnormal areas of railway freight train images according to claim 4, characterized in that: The step of matching and fusing the original image and the initial edited image in the fusion boundary region to obtain an edited image in the abnormal region and realizing regional fusion optimization specifically includes: Texture matching is performed on the original image and the initial edited image in the fusion boundary area, and edge smoothing is performed using a bilateral filter. The calculation formula is: , in, Pixels in the initial edited image for edge smoothing content, is the normalization factor, Pixel Neighborhood, is the pixel point in the initial edited image content, is the pixel in the original image content, Pixel Neighborhood pixels of is the spatial Gaussian filter calculation function, is the pixel Gaussian filter calculation function, pixel point Belongs to the fusion boundary area; Based on the original image, the initial edited image is subjected to illumination correction, and the calculation formula is: , in, The pixels of the initial edited image after lighting correction content, The pixels of the original edited image without light correction content, Calculate function for illumination correction; Based on the original image, color correction is performed on the initial edited image after illumination correction, and the calculation formula is: in, represents the standard deviation on the feature channel, represents the mean value on the feature channel, Compute function for adaptive instance normalization; The corrected initial edited image is fused with the original image to obtain an edited image of the abnormal area.
6. The method for editing abnormal areas of railway freight train images according to claim 5, characterized in that: The corrected initial edited image is fused with the original image to obtain the edited image of the abnormal area. The calculation formula is: , in, Edit the image for abnormal areas, is the original image, is the initial edited image after correction.
7. The method for editing abnormal areas of railway freight train images according to claim 1, characterized in that: The method further comprises: The user interacts by inputting text descriptions, parsing and updating to generate new target semantic information, replacing the original image with the initial edited image, adjusting the target area positioning, and regenerating a new edited image of the abnormal area; Multiple rounds of interaction are performed until a new edited image of the abnormal area that meets user needs is generated.
8. A railway freight train image abnormal area editing system, characterized by: The system is used to implement the method for editing abnormal areas in railway freight train images according to any one of claims 1 to 7, and the system comprises: An editing instruction acquisition module is used to acquire an original image of a railway freight train carriage and editing instructions for abnormal areas; A subtask decomposition module is used to input the editing instructions and the original image into a large language model for parsing, and then gradually decompose the editing instructions into subtask chains through a thought chain reasoning mechanism. The subtask chains include editing object positioning, abnormal content construction, and regional fusion optimization; The editing object positioning module is used to extract target semantic information from the editing instruction, then extract image features from the original image, locate the target area of the image features according to the target semantic information, generate a spatial mask, and output the target area mask to achieve editing object positioning, including: The editing instruction is input into the structured semantic parsing network to obtain a triplet, which includes the editing target object, the editing operation type, and the editing space location, and is expressed as: , in, is a triple, To edit the target object, For edit operation type, To locate the editing space, The triples are input into the pre-trained text editor BERT to extract the target semantic information. The calculation formula is: , in, is the target semantic information, The calculation function for the pre-trained file editor BERT, for dimensional real vector space, The original image is input into the pre-trained SAM editor to extract image features, which are expressed as: , in, is the image feature, is the calculation function of the pre-trained SAM editor, is the original image, for dimensional real vector space, 、 、 are the height, width, and number of channels of the image features, respectively. The target semantic information and the image features are mapped to the same vector space, and the similarity between the image features and the target semantic information is calculated for each spatial position to obtain a heat map. The calculation formula is: , in, For image features Feature information of linear projection on position, is the linear projection matrix of image features, For image features Feature information on the location, for dimensional real vector space, is the semantic feature of the target semantic information on the linear projection, is the linear projection matrix of the target semantic information, Heat map exist The similarity in position, , , The heat map is binarized to obtain a spatial mask, which is calculated as follows: , in, Space mask exist The binary content of the position, Represents a spatial mask The position distribution of the editing target area, 1 represents the editing target area, 0 represents the non-editing target area, is the heatmap threshold, ; An abnormal content construction module is configured to fuse the target semantic information with the image features to obtain a semantic control vector, perform image editing based on the semantic control vector, the spatial mask, and the image features to obtain an initial edited image, and implement abnormal content construction, including: The original image is initialized with noise in the area of the spatial mask to obtain an initial noise image. The calculation formula is as follows: , in, is the pixel point in the initial noise image content, To edit the target area, is the pixel point in the spatial mask content, is the pixel in the original image content, represents the standard normal noise initialization function, The initial noisy image is used as the starting image of the diffusion process. Based on the image features, the semantic control vector and the spatial mask, diffusion denoising is performed step by step to obtain the initial edited image. The calculation formula is as follows: , in, is the time step The predicted denoised image, is the calculation function of the noise prediction model, is the time step The predicted denoised image until =0, get the initial edited image; The regional fusion optimization module is used to pre-process the target area mask to obtain a fusion boundary area, match and fuse the original image with the initial edited image in the fusion boundary area to obtain an abnormal area edited image, and realize regional fusion optimization.
Citation Information
Patent Citations
Text-guided controllable portrait generation method, system and equipment based on diffusion model
CN118114124A
Generalized robot operation method and system based on segmentation mask representation
CN120107583A