Text description generation method and device and electronic equipment
By using multiple cameras on an autonomous driving vehicle to capture the first viewing angle image and synthesize the third viewing angle image, the problem in the prior art that information other than the main vehicle in the video cannot be accurately described, and the accuracy of text description is improved.
Patent Information
- Application Number
- CN202411971712.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art cannot accurately describe information such as vehicles, pedestrians, etc. in the video except for the main vehicle, resulting in a low accuracy of the generated text description.
A plurality of first viewing angle images are captured by a plurality of cameras with different orientations on the target vehicle, a third viewing angle image containing the target vehicle is synthesized based on these images, and inputted to a pre-trained description generation model to output a text description of the target vehicle.
Improve the accuracy of the generated text description and enable more accurate description of the behavior and environmental information of the target vehicle.
Smart Images

Figure CN119992489A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device and electronic device for generating text description. Background Art
[0002] At present, a vehicle with an autonomous driving system is equipped with a camera. When the vehicle with the autonomous driving system is driving as the main vehicle, the camera shoots a video, the video is input into a pre-trained model, and a text description of the video is output. In the related art, the pre-trained model is a general model, and the video after visualization is input into the pre-trained model. The output text description includes a description of the main vehicle's behavior, but it is unable to accurately describe vehicles, pedestrians, and other information in the video other than the main vehicle, resulting in a low accuracy rate of the generated text description. Summary of the invention
[0003] In view of this, an object of the present invention is to provide a method, device and electronic device for generating text descriptions, so as to improve the accuracy of the generated text descriptions.
[0004] In a first aspect, an embodiment of the present invention provides a method for generating a text description, the method comprising: acquiring a plurality of first-perspective images taken of a target vehicle; wherein the plurality of first-perspective images are taken by a plurality of cameras on the target vehicle at the same time; the plurality of cameras have different orientations; the first-perspective images do not include the target vehicle, or include a partial part of the target vehicle; based on the plurality of first-perspective images, synthesizing a third-perspective image including the target vehicle; wherein the third-perspective image includes the complete part of the target vehicle at a specified perspective; inputting the third-perspective image into a pre-trained description generation model, and outputting a text description corresponding to the target vehicle; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of a specified vehicle; the text description includes a behavior description and / or an environmental description of the target vehicle.
[0005] The above-mentioned step of synthesizing a third-perspective image containing a target vehicle based on multiple first-perspective images includes: establishing a three-dimensional scene model based on multiple first-perspective images; wherein the three-dimensional scene model includes the target vehicle and objects in the multiple first-perspective images; generating a virtual camera in the three-dimensional scene model; wherein the virtual camera has a specified positional relationship with the target vehicle; and capturing a third-perspective image containing the target vehicle through the virtual camera.
[0006] Before the above-mentioned step of inputting the third-perspective image into a pre-trained description generation model and outputting a text description corresponding to the target vehicle, the method also includes: identifying objects contained in the third-perspective image and determining the distance between the objects and the target vehicle; determining from the third-perspective image an image area occupied by objects whose distance is greater than a preset distance threshold; deleting the image area from the third-perspective image to obtain a first image; wherein the first image contains objects whose distance is less than or equal to the preset distance threshold; and obtaining a processed third-perspective image based on the first image.
[0007] The step of obtaining the processed third-viewing angle image based on the first image includes: removing color information in the first image to obtain the third-viewing angle image containing line information.
[0008] After the above step of removing the color information in the first image to obtain the third-perspective image containing line information, the method also includes: determining the target color based on the object type of the object contained in the third-perspective image; and filling the image area occupied by the object in the third-perspective image with the target color.
[0009] After the above step of removing the color information in the first image to obtain the third perspective image containing line information, the method also includes: obtaining the movement status data of the target vehicle; wherein the movement status data includes: movement direction and / or movement speed; adding the movement status data to the third perspective image.
[0010] The above-mentioned step of obtaining multiple first-person perspective images taken by the target vehicle includes: obtaining designated index data of the target vehicle; wherein the designated index data includes: at least one of speed, acceleration, steering angle and obstacle distance; the designated index data is associated with the collection time of the designated index data; determining the first time corresponding to the designated index data that meets the preset conditions, and determining the images taken by multiple cameras at the first time as multiple first-person perspective images.
[0011] In a second aspect, an embodiment of the present invention further provides a device for generating a text description, the device comprising: a first-perspective image acquisition module, for acquiring multiple first-perspective images taken of a target vehicle; wherein the multiple first-perspective images are taken by multiple cameras on the target vehicle at the same time; the multiple cameras have different orientations; the first-perspective image does not include the target vehicle, or includes a partial part of the target vehicle; a third-perspective image synthesis module, for synthesizing a third-perspective image containing the target vehicle based on multiple first-perspective images; wherein the third-perspective image includes the complete part of the target vehicle at a specified perspective; a text description output module, for inputting the third-perspective image into a pre-trained description generation model, and outputting a text description corresponding to the target vehicle; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of a specified vehicle; the text description includes a behavior description and / or an environmental description of the target vehicle.
[0012] The embodiments of the present invention bring the following beneficial effects:
[0013] The embodiments of the present invention provide a method, device and electronic device for generating text descriptions, which obtain multiple first-perspective images taken by a target vehicle; wherein the multiple first-perspective images are taken by multiple cameras on the target vehicle at the same time; the multiple cameras have different orientations; the first-perspective images do not include the target vehicle, or include local parts of the target vehicle; based on the multiple first-perspective images, a third-perspective image containing the target vehicle is synthesized; wherein the third-perspective image includes the complete parts of the target vehicle at a specified perspective; the third-perspective image is input into a pre-trained description generation model, and a text description corresponding to the target vehicle is output; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of the specified vehicle; the text description includes a behavior description and / or an environment description of the target vehicle.
[0014] In this method, multiple first-perspective images are taken by multiple cameras in different directions on the target vehicle, a third-perspective image is synthesized based on the multiple first-perspective images, and the third-perspective image is input into a pre-trained model to obtain a text description corresponding to the target vehicle. This method synthesizes multiple first-perspective images to obtain a third-perspective image. The third-perspective image is compatible with the pre-trained model, thereby improving the accuracy of the generated text description.
[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A flowchart of a method for generating a text description provided by an embodiment of the present invention;
[0019] Figure 2 A logic flow chart of a method for generating a text description provided by an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of the structure of a text description generating device provided by an embodiment of the present invention;
[0021] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0023] At present, a vehicle with an autonomous driving system is equipped with a camera. When the vehicle with an autonomous driving system is driving as the main vehicle, the camera shoots video, the video is input into the pre-trained model, and a text description of the video is output, such as 'the main vehicle in the video appears to be a black sedan driving on a multi-lane road', 'as the main vehicle continues to move forward, it changes lanes and turns slightly to the right', etc.
[0024] In the related art, the pre-trained model is a general model. If the video is visualized, the objects in the video are rendered in the visualization process; then the visualized video is input into the pre-trained model, and the output text description includes a description of the main vehicle's behavior, but it is unable to accurately describe vehicles, pedestrians and other information in the video other than the main vehicle, resulting in a low accuracy rate of the generated text description.
[0025] Based on this, a text description generation method, device and electronic device provided by an embodiment of the present invention can be applied to electronic devices such as computers, and in particular, can be applied to devices with a text description generation function, such as generating text descriptions from the first perspective of a vehicle with an autonomous driving system.
[0026] To facilitate understanding of this embodiment, a method for generating a text description disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the method comprises the following steps:
[0027] Step S102, obtaining a plurality of first-perspective images taken by the target vehicle; wherein the plurality of first-perspective images are taken by a plurality of cameras on the target vehicle at the same time; the plurality of cameras are directed in different directions; the first-perspective images do not include the target vehicle, or include a partial part of the target vehicle;
[0028] The target vehicle mentioned above usually refers to a vehicle with an automatic driving system. This embodiment is described with the target vehicle as the main vehicle. The above-mentioned multiple first-person perspective images usually contain environmental information around the target vehicle, such as vehicles, pedestrians, traffic signs, buildings, etc. around the target vehicle. The above-mentioned multiple cameras can be installed in multiple positions of the target vehicle. Taking six cameras as an example, a main camera is installed in front of the target vehicle, a rear camera is installed behind the target vehicle, and cameras are installed in the left front, right front, left rear, and right rear of the target vehicle for shooting. Here, a higher weight value can be set for the initial image taken by the main camera.
[0029] In practical applications, first-person perspective videos can be captured by multiple cameras on the target vehicle, from which specified indicator data that meet the requirements, such as speed, acceleration, pitch angle, etc., can be screened, and multiple first-person perspective images captured by multiple cameras can be obtained at the time corresponding to the specified indicator data that meet the requirements.
[0030] Step S104, synthesizing a third-perspective image including the target vehicle based on the multiple first-perspective images; wherein the third-perspective image includes the complete parts of the target vehicle at a specified perspective;
[0031] The third-view image generally includes the environment information around the target vehicle, the complete part of the target vehicle at a specified viewing angle, the current driving direction and speed information of the target vehicle, etc. The specified viewing angle may be directly above, above left, above right, above rear, etc. of the target vehicle, and the complete part of the target vehicle at the specified viewing angle is displayed.
[0032] Specifically, information in multiple first-perspective images is extracted, and a three-dimensional scene model is established based on the information. In the three-dimensional scene model, the target vehicle in the first-perspective image and the target objects in the multiple first-perspective images are restored, and a virtual camera is generated. According to the positional relationship between the virtual camera and the target vehicle, a third-perspective image including the complete part of the target vehicle at a specified perspective is captured by the virtual camera.
[0033] Step S106, input the third-person perspective image into a pre-trained description generation model, and output a text description corresponding to the target vehicle; wherein the description generation model is pre-trained using sample images, and the sample images are third-person perspective images of the specified vehicle; the text description includes a behavior description and / or an environment description of the target vehicle.
[0034] The above-mentioned description generation model is usually pre-trained, for example, the description generation model is trained by the third-person perspective image of a specified vehicle, and the specified vehicle may be the target vehicle or other vehicles except the target vehicle. For example, the third-person perspective image of the target vehicle is input into the pre-trained description generation model, and the text description corresponding to the target vehicle is output. The text description may be a behavior description of the target vehicle, such as 'the main vehicle is in the left lane, driving at a moderate speed'; or an environmental description of the target vehicle, such as 'there are other vehicles in the lane next to the main vehicle, and there are some traffic signs'; or a behavior description and environmental description of the target vehicle.
[0035] An embodiment of the present invention provides a method for generating a text description, which obtains multiple first-perspective images taken of a target vehicle; wherein the multiple first-perspective images are taken by multiple cameras on the target vehicle at the same time; the multiple cameras have different orientations; the first-perspective images do not include the target vehicle, or include a partial part of the target vehicle; based on the multiple first-perspective images, a third-perspective image containing the target vehicle is synthesized; wherein the third-perspective image includes the complete part of the target vehicle at a specified perspective; the third-perspective image is input into a pre-trained description generation model, and a text description corresponding to the target vehicle is output; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of the specified vehicle; the text description includes a behavior description and / or an environment description of the target vehicle.
[0036] In this method, multiple first-perspective images are taken by multiple cameras in different directions on the target vehicle, a third-perspective image is synthesized based on the multiple first-perspective images, and the third-perspective image is input into a pre-trained model to obtain a text description corresponding to the target vehicle. This method synthesizes multiple first-perspective images to obtain a third-perspective image. The third-perspective image is compatible with the pre-trained model, thereby improving the accuracy of the generated text description.
[0037] In one embodiment, based on multiple first-view images, a third-view image including a target vehicle is synthesized, and the following specific implementation is provided:
[0038] Based on multiple first-perspective images, a three-dimensional scene model is established; wherein the three-dimensional scene model includes a target vehicle and objects in the multiple first-perspective images; a virtual camera is generated in the three-dimensional scene model; wherein the virtual camera has a specified positional relationship with the target vehicle; and a third-perspective image including the target vehicle is captured by the virtual camera.
[0039] The above-mentioned three-dimensional scene model can be specifically established by laser point cloud scanning of real scene data, multiple first-person perspective images as auxiliary data, remote sensing data, etc. The objects in the above-mentioned multiple first-person perspective images can usually be one or more, such as pedestrians, obstacles, other vehicles other than the target vehicle, etc. The above-mentioned virtual camera has a specified positional relationship with the target vehicle, such as the virtual camera is located at the upper rear, lower right, etc. of the target vehicle.
[0040] In actual implementation, a three-dimensional scene model can be established through the relative positional relationship of objects in multiple perspective images and remote sensing data. The target vehicle and objects in multiple first-perspective images can be displayed in the three-dimensional scene model, and a virtual camera can be generated at a spatial coordinate point with a specified positional relationship with the target vehicle. Then, a third-perspective image containing the target vehicle can be captured by the virtual camera.
[0041] Before the above step of inputting the third-view image into the pre-trained description generation model and outputting the text description corresponding to the target vehicle, the following optional specific implementation methods are provided:
[0042] In an optional method, an object contained in a third-perspective image is identified, and the distance between the object and a target vehicle is determined; an image area occupied by an object whose distance is greater than a preset distance threshold is determined from the third-perspective image; the image area is deleted from the third-perspective image to obtain a first image; wherein the first image contains an object whose distance is less than or equal to the preset distance threshold; and a processed third-perspective image is obtained based on the first image.
[0043] The preset distance threshold is usually the maximum distance between the object in the third-view image and the target vehicle. The image area usually refers to the pixel position occupied by the object in the third-view image, and can be a square, circular, sector-shaped area, etc.
[0044] In actual implementation, the object contained in the third-perspective image is identified, the real distance between the object and the target vehicle is calculated, and the image area occupied by the object whose distance is greater than a preset distance threshold is calculated from the third-perspective image. The image area is then deleted from the third-perspective image, and the deleted image area is merged with the color of the surrounding area to obtain the first image. The first image is further processed, such as using a VIZ rendering tool, to obtain a processed third-perspective image.
[0045] Furthermore, the color information in the first image is removed to obtain a third-viewing angle image containing line information.
[0046] The above color information generally refers to the color of a real object, specifically, the RGB pixel value in the first image.
[0047] For example, all color information within the outline of the object in the first image is removed, and all color information outside the outline of the object in the first image is removed to obtain a third-view image containing line information.
[0048] Furthermore, based on the object type of the object contained in the third-view image, a target color is determined; and the image area occupied by the object in the third-view image is filled with the target color.
[0049] The above object types may specifically include a host vehicle, a road, a lane dividing line, a pedestrian, other vehicles except the host vehicle, etc. The above target color may generally be white, dark gray, fluorescent powder, etc.
[0050] For example, if the object types of the objects included in the third viewing angle image are roads, other vehicles and pedestrians, the image areas occupied by the objects in the third image may be filled with dark gray, white and fluorescent powder respectively.
[0051] Furthermore, the movement status data of the target vehicle is obtained; wherein the movement status data includes: movement direction and / or movement speed; and the movement status data is added to the third-view image.
[0052] For example, the movement state data of the target vehicle corresponding to the third-view image is obtained, and the movement state data may be the movement direction, the movement speed, or the movement direction and the movement speed; and then the movement state data is added to the third-view image.
[0053] In one embodiment, the step of obtaining multiple first-person perspective images taken by the target vehicle provides the following specific implementation method:
[0054] Obtain designated index data of a target vehicle; wherein the designated index data includes: at least one of speed, acceleration, steering angle and obstacle distance; the designated index data is associated with a collection time of the designated index data; determine a first time corresponding to the designated index data that meets a preset condition, and determine images captured by multiple cameras at the first time as multiple first-perspective images.
[0055] The above-mentioned preset conditions can usually be to set a weight value for at least one specified indicator data, such as acceleration has the highest weight and steering angle has the lowest weight, sort at least one specified indicator data according to the weight value, select the top-ranked specified indicator data and determine the first time. The above-mentioned speed, acceleration, and steering angle can be detected by sensors, and the obstacle distance can be calculated by point cloud data.
[0056] It is understandable that in some scenarios, the distance to obstacles will be shortened due to the change in acceleration and steering angle when the target vehicle is turning, meeting other vehicles, or braking suddenly. Therefore, these specified indicator data can be used to assist in judging the first time and then determine the first-person perspective image.
[0057] In actual implementation, at least one specified index data of the target vehicle among speed, acceleration, steering angle and obstacle distance is obtained, and a weight value is set for at least one specified index data; at least one specified index data is sorted according to the weight value, and the collection time associated with the top-ranked specified index data is determined as the first time, and the images taken by multiple cameras at the first time are determined as multiple first-perspective images.
[0058] exist Figure 2 In the text description generation method, the logical flow is provided, which is described in detail below.
[0059] Step S201, capturing a first-perspective image; specifically, capturing multiple first-perspective images by multiple cameras installed on the target vehicle;
[0060] Step S202, adding sensors to collect designated indicator data; specifically, adding sensors to the target vehicle to collect speed, acceleration, and steering angle, and calculating the obstacle distance through point cloud data.
[0061] Step S203, rendering the image by using a VIZ rendering tool; specifically, rendering the first-person perspective image by using the VIZ rendering tool;
[0062] Step S204, removing noise from the image; specifically, deleting the image area occupied by the object whose distance from the target vehicle exceeds a preset distance threshold from the third-view image;
[0063] Step S205, inputting the image into a description generation model; specifically, inputting the third-view image into a pre-trained description generation model;
[0064] Step S206, generating a text description corresponding to the target vehicle; specifically, the description generation model generates a text description corresponding to the target vehicle from the first perspective.
[0065] Corresponding to the above method embodiment, the present invention provides a device for generating text description, such as Figure 3 As shown, the device comprises:
[0066] The first perspective image acquisition module 31 is used to acquire a plurality of first perspective images taken by the target vehicle; wherein the plurality of first perspective images are taken by a plurality of cameras on the target vehicle at the same time; the plurality of cameras are directed in different directions; the first perspective images do not include the target vehicle, or include a partial part of the target vehicle;
[0067] A third-perspective image synthesis module 32 is used to synthesize a third-perspective image containing the target vehicle based on the multiple first-perspective images; wherein the third-perspective image includes the complete parts of the target vehicle at a specified perspective;
[0068] The text description output module 33 is used to input the third-person perspective image into a pre-trained description generation model and output a text description corresponding to the target vehicle; wherein the description generation model is pre-trained using sample images, and the sample images are third-person perspective images of the specified vehicle; the text description includes a behavior description and / or an environment description of the target vehicle.
[0069] An embodiment of the present invention provides a device for generating a text description, which obtains a plurality of first-perspective images taken of a target vehicle; wherein the plurality of first-perspective images are taken by a plurality of cameras on the target vehicle at the same time; the plurality of cameras have different orientations; the first-perspective images do not include the target vehicle, or include a partial part of the target vehicle; based on the plurality of first-perspective images, a third-perspective image including the target vehicle is synthesized; wherein the third-perspective image includes the complete part of the target vehicle at a specified perspective; the third-perspective image is input into a pre-trained description generation model, and a text description corresponding to the target vehicle is output; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of a specified vehicle; the text description includes a behavior description and / or an environment description of the target vehicle.
[0070] In this method, multiple first-perspective images are taken by multiple cameras in different directions on the target vehicle, a third-perspective image is synthesized based on the multiple first-perspective images, and the third-perspective image is input into a pre-trained model to obtain a text description corresponding to the target vehicle. This method synthesizes multiple first-perspective images to obtain a third-perspective image. The third-perspective image is compatible with the pre-trained model, thereby improving the accuracy of the generated text description.
[0071] The above-mentioned third-perspective image synthesis module is also used to establish a three-dimensional scene model based on multiple first-perspective images; wherein the three-dimensional scene model includes a target vehicle and objects in multiple first-perspective images; generate a virtual camera in the three-dimensional scene model; wherein the virtual camera has a specified positional relationship with the target vehicle; and capture a third-perspective image containing the target vehicle through the virtual camera.
[0072] The above-mentioned device also includes a first image acquisition module, which is used to identify the object contained in the third-perspective image and determine the distance between the object and the target vehicle; determine the image area occupied by the object whose distance is greater than a preset distance threshold from the third-perspective image; delete the image area from the third-perspective image to obtain a first image; wherein the first image contains an object whose distance is less than or equal to the preset distance threshold; and obtain a processed third-perspective image based on the first image.
[0073] The first image acquisition module is further used to remove color information from the first image to obtain a third viewing angle image containing line information.
[0074] The above device also includes a target color filling module, which is used to determine the target color based on the object type contained in the third-view image; and fill the image area occupied by the object in the third-view image with the target color.
[0075] The above-mentioned device also includes a movement state data acquisition module, which is used to acquire the movement state data of the target vehicle; wherein the movement state data includes: movement direction and / or movement speed; and the movement state data is added to the third-view image.
[0076] The above-mentioned first-perspective image acquisition module is also used to: obtain designated index data of the target vehicle; wherein the designated index data includes: at least one of speed, acceleration, steering angle and obstacle distance; the designated index data is associated with the collection time of the designated index data; determine the first time corresponding to the designated index data that meets the preset conditions, and determine the images taken by multiple cameras at the first time as multiple first-perspective images.
[0077] This embodiment also provides an electronic device, including a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the above-mentioned text description generation method. The electronic device can be a server or a terminal device.
[0078] See also Figure 4 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer executable instructions that can be executed by the processor 100. The processor 100 executes the computer executable instructions to implement the above-mentioned text description generation method.
[0079] Further, Figure 4 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .
[0080] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0081] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 100. The above processor 100 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and completes the steps of the method of the above embodiment in combination with its hardware.
[0082] This embodiment also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned text description generation method.
[0083] The computer program product of the text description generation method, device and electronic device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiments. The specific implementation can be found in the method embodiments and will not be repeated here.
[0084] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0085] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0086] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0087] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0088] Finally, it should be noted that the above embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for generating a text description, characterized in that: The method comprises: Acquire multiple first-perspective images taken by the target vehicle; wherein the multiple first-perspective images are taken by multiple cameras on the target vehicle at the same time; the multiple cameras are facing different directions; the first-perspective images do not include the target vehicle, or include a partial part of the target vehicle; Based on the multiple first-perspective images, synthesizing a third-perspective image containing the target vehicle; wherein the third-perspective image includes the complete part of the target vehicle at a specified perspective; The third-perspective image is input into a pre-trained description generation model, and a text description corresponding to the target vehicle is output; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of a specified vehicle; and the text description includes a behavior description and / or an environment description of the target vehicle.
2. The method according to claim 1, characterized in that The step of synthesizing a third perspective image including the target vehicle based on the multiple first perspective images comprises: Based on the multiple first-perspective images, a three-dimensional scene model is established; wherein the three-dimensional scene model includes the target vehicle and the objects in the multiple first-perspective images; Generate a virtual camera in the three-dimensional scene model; wherein the virtual camera has a specified positional relationship with the target vehicle; A third-perspective image including the target vehicle is captured by the virtual camera.
3. The method according to claim 1, characterized in that Before the step of inputting the third-view image into a pre-trained description generation model to output a text description corresponding to the target vehicle, the method further includes: Identify an object contained in the third-view image, and determine a distance between the object and the target vehicle; Determine, from the third viewing angle image, an image area occupied by the object whose distance is greater than a preset distance threshold; Deleting the image area from the third-view image to obtain a first image; wherein the first image contains an object whose distance is less than or equal to the preset distance threshold; The processed third-viewing angle image is obtained based on the first image.
4. The method according to claim 3, characterized in that The step of obtaining the processed third-viewing angle image based on the first image comprises: The color information in the first image is removed to obtain a third-viewing angle image containing line information.
5. The method according to claim 4, characterized in that After removing the color information in the first image to obtain the third viewing angle image containing line information, the method further includes: determining a target color based on an object type of the object included in the third viewing angle image; An image area occupied by the object in the third viewing angle image is filled with the target color.
6. The method according to claim 5, characterized in that After removing the color information in the first image to obtain the third viewing angle image containing line information, the method further includes: Acquire the movement state data of the target vehicle; wherein the movement state data includes: movement direction and / or movement speed; The movement status data is added to the third-perspective image.
7. The method according to claim 1, characterized in that The step of acquiring a plurality of first-person perspective images taken by a target vehicle comprises: Acquire designated index data of the target vehicle; wherein the designated index data includes: at least one of speed, acceleration, steering angle and obstacle distance; the designated index data is associated with the collection time of the designated index data; A first time corresponding to designated indicator data satisfying a preset condition is determined, and images captured by a plurality of cameras at the first time are determined as a plurality of first-viewing-angle images.
8. A device for generating text description, characterized in that: The device comprises: A first-perspective image acquisition module is used to acquire a plurality of first-perspective images taken by a target vehicle; wherein the plurality of first-perspective images are taken by a plurality of cameras on the target vehicle at the same time; the plurality of cameras are facing different directions; the first-perspective images do not include the target vehicle, or include a partial part of the target vehicle; A third-perspective image synthesis module, used to synthesize a third-perspective image containing the target vehicle based on the multiple first-perspective images; wherein the third-perspective image includes the complete part of the target vehicle at a specified perspective; A text description output module is used to input the third-perspective image into a pre-trained description generation model, and output a text description corresponding to the target vehicle; wherein the description generation model is pre-trained using sample images, and the sample images are third-perspective images of a specified vehicle; and the text description includes a behavior description and / or an environment description of the target vehicle.
9. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method for generating a text description according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method for generating a text description according to any one of claims 1 to 7.
Citation Information
Cited By
Vehicle-mounted terminal data processing method and device based on YTS engine and electronic equipment
CN120151590A
Vehicle terminal data processing method, device and electronic equipment based on YTS engine
CN120151590B