Model training method and device and electronic equipment

By resolving information and text labels to the initial image, the training and description generation model is solved, and the existing pre-trained model has low accuracy when identifying the behavior of the main car, achieving higher recognition accuracy and better adaptability to complex environments.

CN119992490APending Publication Date: 2025-05-13文远京行(北京)科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411971714.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing pre-trained models have identification errors when identifying the behavior of the main car, resulting in a low accuracy rate of recognition of the behavior of the main car when the traffic volume is large and the environment is complex.

Method used

By performing information removal processing on the initial image, an intermediate image containing the specified image information is obtained, and the intermediate image is associated with the text label corresponding to the initial image. The description generation model is trained based on the intermediate image and the text label, and the text description of the target vehicle is output.

Benefits of technology

It improves the accuracy of the model's recognition of the main car behavior, reduces the impact of noise data, and enhances the recognition ability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992490A_ABST
    Figure CN119992490A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device and electronic equipment, and the method comprises the steps: obtaining an initial image and a text tag corresponding to the initial image; wherein the text label comprises a text description of the initial image; the text description comprises behavior description and / or environment description of the target vehicle; performing information removal processing on the initial image to obtain an intermediate image containing specified image information, and associating the intermediate image with a text label corresponding to the initial image; wherein the intermediate image at least comprises line information of at least part of objects in the initial image; and training a description generation model based on the intermediate image and the text tag corresponding to the intermediate image to obtain a trained description generation model. In the mode, the model is trained through the processed image and the corresponding text tag, so that the accuracy of the model for identifying the behavior of the main vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model training method, device and electronic equipment. Background Art

[0002] At present, a vehicle with an automatic driving system is equipped with a camera. When the vehicle with the automatic driving system is driving as the main vehicle, the camera shoots a video, the video is input into a pre-trained model, and a text description of the video is output. In the related art, the pre-trained model is a general model. After the video is input into the pre-trained model, the output text description contains a description of the main vehicle's behavior, but also contains noise data, such as "the sun is shining, indicating that it is a sunny day", "trees are on both sides of the road", etc. When the traffic volume is large and the environment is complex, the pre-trained model may make recognition errors, resulting in a low accuracy rate in recognizing the main vehicle's behavior. Summary of the invention

[0003] In view of this, an object of the present invention is to provide a model training method, device and electronic equipment to improve the accuracy of host vehicle behavior recognition.

[0004] In a first aspect, an embodiment of the present invention provides a model training method, the method comprising: obtaining an initial image, and a text label corresponding to the initial image; wherein the text label comprises a text description of the initial image; the text description comprises a behavior description and / or an environment description of a target vehicle; performing information removal processing on the initial image to obtain an intermediate image containing specified image information, and associating the intermediate image with the text label corresponding to the initial image; wherein the intermediate image comprises at least: line information of at least part of the object in the initial image; training a description generation model based on the intermediate image and the text label corresponding to the intermediate image to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model to output a text description corresponding to the target vehicle.

[0005] The above-mentioned step of performing information removal processing on the initial image to obtain an intermediate image containing specified image information includes: identifying objects contained in the initial image and determining the distance between the objects and the target vehicle that photographed the initial image; determining from the initial image the image area occupied by objects whose distance is greater than a preset distance threshold; deleting the image area from the initial image to obtain a first image; wherein the first image contains objects whose distance is less than or equal to the preset distance threshold; and generating an intermediate image based on the first image.

[0006] The above-mentioned step of generating an intermediate image based on the first image includes: if the specified vehicle does not exist in the first image, generating the specified vehicle in the first image and adding a text description of the specified vehicle in the text label; removing the color information in the first image to obtain an intermediate image containing line information.

[0007] After the above step of removing the color information in the first image to obtain the intermediate image containing line information, the method further includes: determining the target color based on the object type of the object contained in the intermediate image; and filling the image area occupied by the object in the intermediate image with the target color.

[0008] After the above step of removing the color information in the first image to obtain the intermediate image containing line information, the above method also includes: obtaining the movement status data of the target vehicle that took the initial image; wherein the movement status data includes: moving direction and / or moving speed; adding the movement status data to the intermediate image.

[0009] The above-mentioned initial image includes multiple images, and the multiple initial images are taken by multiple cameras on the same target vehicle at the same time; the multiple cameras have different directions.

[0010] The above-mentioned step of obtaining the initial image includes: obtaining a video shot by a camera on the target vehicle, and specified index data when the target vehicle shoots the video; wherein the video includes multiple video frames; the specified index data includes: at least one of speed, acceleration, steering angle and obstacle distance; based on time parameters, determining the specified index data corresponding to each video frame; and determining the video frame whose specified index data meets preset conditions as the initial image.

[0011] In a second aspect, an embodiment of the present invention further provides a model training device, which includes: an image and label acquisition module, used to acquire an initial image, and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of the target vehicle; an information removal processing module, used to perform information removal processing on the initial image to obtain an intermediate image containing specified image information, and associate the intermediate image with the text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image; a model training module, used to train a description generation model based on the intermediate image and the text label corresponding to the intermediate image, to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model, and a text description corresponding to the target vehicle is output.

[0012] The embodiments of the present invention bring the following beneficial effects:

[0013] The embodiments of the present invention provide a model training method, device and electronic device, which obtain an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of a target vehicle; information removal processing is performed on the initial image to obtain an intermediate image containing specified image information, and the intermediate image is associated with the text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image; based on the intermediate image and the text label corresponding to the intermediate image, a description generation model is trained to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model, and a text description corresponding to the target vehicle is output.

[0014] In this method, an intermediate image is obtained by removing information from the initial image, and the model is trained based on the intermediate image and the corresponding text labels to obtain a trained model. This method trains the model with processed images and corresponding text labels, which is beneficial to improving the accuracy of the model in recognizing the main vehicle behavior.

[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flowchart of a model training method provided by an embodiment of the present invention;

[0019] Figure 2 A logic flow chart of a model training method provided by an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of the structure of a model training device provided by an embodiment of the present invention;

[0021] Figure 4A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0023] At present, a vehicle with an autonomous driving system is equipped with a camera. When the vehicle with the autonomous driving system is driving as the main vehicle, the camera is used to shoot video, the video is input into a pre-trained model, and a text description of the video is output.

[0024] In the related art, the pre-trained model is a general model. After the video is input into the pre-trained model, the output text description contains a description of the main vehicle's behavior, but also contains noise data, such as "the sun is shining, indicating that it is a sunny day", "trees are on both sides of the road", etc. When the traffic volume is large and the environment is complex, the pre-trained model may make recognition errors, resulting in a low accuracy rate in identifying the main vehicle's behavior.

[0025] Based on this, a model training method, device and electronic device provided in an embodiment of the present invention can be applied to electronic devices such as computers, and in particular, can be applied to devices with text description generation functions, such as training text description generation models.

[0026] To facilitate understanding of this embodiment, a model training method disclosed in an embodiment of the present invention is first introduced in detail. Figure 1 As shown, the method comprises the following steps:

[0027] Step S102, obtaining an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of the target vehicle;

[0028] The above-mentioned initial images are usually multiple, and the multiple initial images are taken by multiple cameras installed on the target vehicle at the same time. The above-mentioned target vehicle usually refers to a vehicle with an automatic driving system. This embodiment is described with the target vehicle as the main vehicle.

[0029] In practical applications, an initial video can be captured by multiple cameras on a target vehicle, from which specified indicator data that meet the requirements, such as speed, acceleration, pitch angle, etc., are screened, and a video frame is determined according to the specified indicator data that meet the requirements, and the video frame is obtained as an initial image; and a text label corresponding to the initial image is obtained, and the text label includes a text description of the initial image, which can be a behavior description of the target vehicle, such as 'the main vehicle is located in the left lane and travels at a moderate speed'; or a description of the environment of the target vehicle, such as 'there are other vehicles and some traffic signs in the lane next to the main vehicle'; or a description of the behavior and environment of the target vehicle.

[0030] Step S104, performing information removal processing on the initial image to obtain an intermediate image containing the specified image information, and associating the intermediate image with the text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image;

[0031] The above information removal process usually refers to deleting the image area where the object meets the preset conditions is located, or removing the color information in the image. The above specified image information can usually be the line information of the object, or the image area filled with the target color, or the moving state data of the target vehicle.

[0032] Specifically, objects contained in the initial image are identified, image areas occupied by objects whose distance to the target vehicle is greater than a preset distance threshold are determined, and the image areas are deleted from the initial image to obtain a first image, and color information in the first image is removed to obtain an intermediate image containing line information of one or more objects in the initial image; further, the objects contained in the intermediate image can be filled with a target color.

[0033] Step S106, based on the intermediate image and the text label corresponding to the intermediate image, the description generation model is trained to obtain a trained description generation model; wherein, after the intermediate image is input into the trained description generation model, the text description corresponding to the target vehicle is output.

[0034] In practical applications, the intermediate image and the text label corresponding to the intermediate image are input into the description generation model to train the description generation model. For example, if there are no other vehicles except the target vehicle in the intermediate image, then other vehicles can be rendered in the intermediate image and text descriptions can be added to the text labels. Then, from the first-person perspective of the target vehicle, the rendered intermediate image and the text label corresponding to the rendered intermediate image are input into the description generation model.

[0035] After obtaining the trained description generation model, the intermediate image is input into the trained description generation model to obtain a text description corresponding to the target vehicle.

[0036] An embodiment of the present invention provides a model training method, which obtains an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of a target vehicle; information removal processing is performed on the initial image to obtain an intermediate image containing specified image information, and the intermediate image is associated with the text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image; based on the intermediate image and the text label corresponding to the intermediate image, a description generation model is trained to obtain a trained description generation model; wherein after the intermediate image is input into the trained description generation model, a text description corresponding to the target vehicle is output.

[0037] In this method, an intermediate image is obtained by removing information from the initial image, and the model is trained based on the intermediate image and the corresponding text labels to obtain a trained model. This method trains the model with processed images and corresponding text labels, which is beneficial to improving the accuracy of the model in recognizing the main vehicle behavior.

[0038] The above step of performing information removal processing on the initial image to obtain an intermediate image containing the specified image information provides the following optional specific implementation methods:

[0039] In an optional method, an object contained in an initial image is identified, and a distance between the object and a target vehicle that photographed the initial image is determined; an image area occupied by an object whose distance is greater than a preset distance threshold is determined from the initial image; the image area is deleted from the initial image to obtain a first image; wherein the first image contains an object whose distance is less than or equal to the preset distance threshold; and an intermediate image is generated based on the first image.

[0040] The preset distance threshold is usually the maximum distance between the object in the initial image and the target vehicle. The image area usually refers to the pixel position occupied by the object in the initial image, and can be a square, circular, sector-shaped area, etc.

[0041] In actual implementation, the objects contained in the initial image are identified, the real distance between the object and the target vehicle is calculated, and the image area occupied by the object whose distance is greater than the preset distance threshold is calculated from the initial image. The image area is then deleted from the initial image, and the deleted image area is merged with the color of the surrounding area to obtain a first image. The first image can be further processed to generate an intermediate image.

[0042] Furthermore, if the designated vehicle does not exist in the first image, the designated vehicle is generated in the first image, and a text description of the designated vehicle is added to the text label; and the color information in the first image is removed to obtain an intermediate image containing line information.

[0043] The color information generally refers to the color of a real object, specifically, the RGB pixel value in the first image. The designated vehicle generally refers to other vehicles except the target vehicle.

[0044] For example, if the specified vehicle does not exist in the intermediate image, the specified vehicle can be rendered in the intermediate image and a text description can be added to the text label.

[0045] Furthermore, all color information within the outline of the object in the first image is removed, and all color information outside the outline of the object in the first image is removed, to obtain an intermediate image containing line information.

[0046] Further, based on the object type of the object contained in the intermediate image, a target color is determined; and the image area occupied by the object in the intermediate image is filled with the target color.

[0047] The above object types may specifically include a host vehicle, a road, a lane dividing line, a pedestrian, other vehicles except the host vehicle, etc. The above target color may generally be white, dark gray, fluorescent powder, etc.

[0048] For example, if the object types of the objects included in the intermediate image are roads, other vehicles, and pedestrians, then the image areas occupied by the objects in the intermediate image may be filled with dark gray, white, and fluorescent powder, respectively.

[0049] Furthermore, the movement state data of the target vehicle that captured the initial image is obtained; wherein the movement state data includes: movement direction and / or movement speed; and the movement state data is added to the intermediate image.

[0050] For example, the movement state data of the target vehicle is obtained when the camera captures the initial image. The movement state data may be the movement direction, the movement speed, or the movement direction and the movement speed. The movement state data is then added to the third-view image.

[0051] In one embodiment, the initial image includes a plurality of images, and the plurality of initial images are captured by a plurality of cameras on the same target vehicle at the same time; and the plurality of cameras are directed in different directions.

[0052] The above-mentioned multiple cameras can be installed in multiple positions of the target vehicle. Taking six cameras as an example, a main camera is installed in front of the target vehicle, a rear camera is installed behind the target vehicle, and cameras are installed at the left front, right front, left rear and right rear of the target vehicle for shooting. Here, a higher weight value can be set for the initial image taken by the main camera.

[0053] The above step of obtaining the initial image provides the following specific implementation methods:

[0054] Obtain a video captured by a camera on a target vehicle, and designated index data when the target vehicle captured the video; wherein the video includes multiple video frames; the designated index data includes at least one of speed, acceleration, steering angle, and obstacle distance; based on a time parameter, determine the designated index data corresponding to each video frame; and determine the video frame whose designated index data meets preset conditions as an initial image.

[0055] The above speed, acceleration and steering angle can be detected by sensors, and the obstacle distance can be calculated by point cloud data.

[0056] It is understandable that in some scenarios, the distance to obstacles will be shortened due to the change in acceleration and steering angle when the target vehicle is turning, meeting other vehicles, or braking suddenly. Therefore, these specified indicator data can be used to assist in judging the first time and then determine the first-person perspective image.

[0057] In actual implementation, the video captured by the camera on the target vehicle and at least one specified indicator data of the target vehicle among speed, acceleration, steering angle and obstacle distance are obtained, and a weight value is set for at least one specified indicator data; under the time parameter, each video frame in the video is sorted by the weight value, and the difference value of the specified indicator data between the current video frame and the previous video frame is compared during sorting, and the final sorting is determined according to the difference value, and the specified number of video frames at the top of the final sorting are determined as the initial images.

[0058] exist Figure 2 In the figure, the logical flow of the model training method is provided, which is described in detail below.

[0059] Step S201, capturing an initial image; specifically, capturing the image using multiple cameras installed on the target vehicle, and selecting video frames from the videos captured by the cameras as the initial image;

[0060] Step S202, rendering the image by using a VIZ rendering tool; specifically, rendering the initial image by using the VIZ rendering tool;

[0061] Step S203, removing noise from the image; specifically, deleting from the initial image the image area occupied by objects whose distance from the target vehicle exceeds a preset distance threshold;

[0062] Step S204, matching the image to the text description in the label and training the description generation model; specifically, matching the intermediate image to the text description in the text label, and then inputting the image into the description generation model for training;

[0063] Step S205, obtaining a description generation model that has completed training.

[0064] Corresponding to the above method embodiment, the present invention provides a model training device, such as Figure 3 As shown, the device comprises:

[0065] The image and label acquisition module 31 is used to acquire an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of the target vehicle;

[0066] The information removal processing module 32 is used to perform information removal processing on the initial image to obtain an intermediate image containing the specified image information, and associate the intermediate image with the text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image;

[0067] The model training module 33 is used to train the description generation model based on the intermediate image and the text label corresponding to the intermediate image to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model, and the text description corresponding to the target vehicle is output.

[0068] An embodiment of the present invention provides a model training device, which obtains an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of a target vehicle; performs information removal processing on the initial image to obtain an intermediate image containing specified image information, and associates the intermediate image with the text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image; based on the intermediate image and the text label corresponding to the intermediate image, trains a description generation model to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model, and outputs a text description corresponding to the target vehicle.

[0069] In this method, an intermediate image is obtained by removing information from the initial image, and the model is trained based on the intermediate image and the corresponding text labels to obtain a trained model. This method trains the model with processed images and corresponding text labels, which is beneficial to improving the accuracy of the model in recognizing the main vehicle behavior.

[0070] The above-mentioned information removal processing module is also used to identify objects contained in the initial image and determine the distance between the objects and the target vehicle that took the initial image; determine the image area occupied by objects whose distance is greater than a preset distance threshold from the initial image; delete the image area from the initial image to obtain a first image; wherein the first image contains objects whose distance is less than or equal to the preset distance threshold; and generate an intermediate image based on the first image.

[0071] The above-mentioned information removal processing module is also used to generate a specified vehicle in the first image if the specified vehicle does not exist in the first image, and add a text description of the specified vehicle in the text label; remove the color information in the first image to obtain an intermediate image containing line information.

[0072] The above-mentioned device also includes a target color filling module, which is used to determine the target color based on the object type contained in the intermediate image; and fill the image area occupied by the object in the intermediate image with the target color.

[0073] The above-mentioned device also includes a movement state data adding module, which is used to obtain the movement state data of the target vehicle that captured the initial image; wherein the movement state data includes: movement direction and / or movement speed; and add the movement state data to the intermediate image.

[0074] The above-mentioned initial image includes multiple images, and the multiple initial images are taken by multiple cameras on the same target vehicle at the same time; the multiple cameras have different directions.

[0075] The above-mentioned image and label acquisition module is also used to obtain the video taken by the camera on the target vehicle, and the specified indicator data when the target vehicle takes the video; wherein the video includes multiple video frames; the specified indicator data includes: at least one of speed, acceleration, steering angle and obstacle distance; based on the time parameter, determine the specified indicator data corresponding to each video frame; and determine the video frame whose specified indicator data meets the preset conditions as the initial image.

[0076] This embodiment also provides an electronic device, including a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the training method of the above model. The electronic device can be a server or a terminal device.

[0077] See also Figure 4 As shown, the electronic device includes a processor 100 and a memory 101, wherein the memory 101 stores computer executable instructions that can be executed by the processor 100, and the processor 100 executes the computer executable instructions to implement the above-mentioned model training method.

[0078] Further, Figure 4 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .

[0079] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0080] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 100. The above processor 100 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and completes the steps of the method of the above embodiment in combination with its hardware.

[0081] This embodiment also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the training method of the above-mentioned model.

[0082] The computer program product of the model training method, device and electronic device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be found in the method embodiments, which will not be repeated here.

[0083] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0084] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0085] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0086] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0087] Finally, it should be noted that the above embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A model training method, characterized in that: The method comprises: Acquire an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of the target vehicle; Performing information removal processing on the initial image to obtain an intermediate image containing designated image information, and associating the intermediate image with a text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image; Based on the intermediate image and the text label corresponding to the intermediate image, a description generation model is trained to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model, and a text description corresponding to the target vehicle is output.

2. The method according to claim 1, characterized in that The step of performing information removal processing on the initial image to obtain an intermediate image containing designated image information comprises: identifying an object contained in the initial image and determining a distance between the object and a target vehicle that captured the initial image; Determine from the initial image the image region occupied by the object whose distance is greater than a preset distance threshold; Deleting the image area from the initial image to obtain a first image; wherein the first image contains an object whose distance is less than or equal to the preset distance threshold; The intermediate image is generated based on the first image.

3. The method according to claim 2, characterized in that The step of generating the intermediate image based on the first image comprises: If the designated vehicle does not exist in the first image, generating the designated vehicle in the first image, and adding a text description of the designated vehicle to the text label; The color information in the first image is removed to obtain an intermediate image containing the line information.

4. The method according to claim 3, characterized in that After removing the color information in the first image to obtain the intermediate image containing the line information, the method further includes: determining a target color based on an object type of an object contained in the intermediate image; The image area occupied by the object in the intermediate image is filled with the target color.

5. The method according to claim 3, characterized in that: After removing the color information in the first image to obtain the intermediate image containing the line information, the method further includes: Acquire movement state data of the target vehicle that captured the initial image; wherein the movement state data includes: movement direction and / or movement speed; The movement status data is added to the intermediate image.

6. The method according to claim 1, characterized in that The initial images include multiple images, and the multiple initial images are captured by multiple cameras on the same target vehicle at the same time; the multiple cameras are facing different directions.

7. The method according to claim 1, characterized in that The steps of obtaining the initial image include: Acquire a video captured by a camera on a target vehicle, and designated index data when the target vehicle captured the video; wherein the video includes a plurality of video frames; and the designated index data includes at least one of speed, acceleration, steering angle, and obstacle distance; Based on the time parameter, determining the specified indicator data corresponding to each of the video frames; The video frame whose specified indicator data meets the preset conditions is determined as the initial image.

8. A model training device, characterized in that: The device comprises: An image and label acquisition module, used to acquire an initial image and a text label corresponding to the initial image; wherein the text label includes a text description of the initial image; the text description includes a behavior description and / or an environment description of the target vehicle; An information removal processing module is used to perform information removal processing on the initial image to obtain an intermediate image containing designated image information, and associate the intermediate image with a text label corresponding to the initial image; wherein the intermediate image at least includes: line information of at least part of the object in the initial image; A model training module is used to train a description generation model based on the intermediate image and the text label corresponding to the intermediate image to obtain a trained description generation model; wherein the intermediate image is input into the trained description generation model, and the text description corresponding to the target vehicle is output.

9. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the training method of the model described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the training method of the model described in any one of claims 1-7.

Citation Information

Cited By

  • Vehicle-mounted terminal data processing method and device based on YTS engine and electronic equipment

    CN120151590A