A video image processing method, device, equipment and storage medium

By using target detection and classification, non-critical targets are identified and replaced with smaller substitutes, solving the problem of inaccurate information transmission in video transmission and achieving efficient video data transmission and reduced latency.

CN116601950BActive Publication Date: 2026-08-25ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180083386.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2026-08-25
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

During the remote transmission of multi-channel video data, existing technical solutions convert the original target objects in the video footage into animations or color blocks, resulting in inaccurate information transmission and increasing driving risks.

Method used

By detecting and classifying objects, non-critical objects are identified and replaced with smaller substitutes, while critical objects are retained, thus reducing video data volume and improving transmission efficiency.

Benefits of technology

It enables the accurate transmission of important information within a short time delay, reduces video transmission latency and data volume, while ensuring video output quality and adapting to network communication limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116601950B_ABST
    Figure CN116601950B_ABST
Patent Text Reader

Abstract

A video image processing method, device, equipment and storage medium, the method comprises: acquiring a first target frame video image, the first target frame video image is any frame video image in a plurality of frames of video image to be processed (S101); target detection is carried out on the first target frame video image, and at least one target object in the first target frame video image is determined (S103); based on the preset target object classification rule, determine the first target object to be processed in the at least one target object (S105); the first target object to be processed is replaced by a preset target substitute in the first target frame video image, and a second target frame video image is obtained (S107); wherein the data amount of the preset target substitute is less than the data amount of the first target object to be processed. By using the method, the data amount of the video can be reduced without affecting the actual output effect, the transmission rate of the video is improved, and the transmission delay of the video is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to a video image processing method, apparatus, device, and storage medium. Background Technology

[0002] Currently, industries such as security, healthcare, and automotive are all relying on visual perception or visual monitoring for the transmission and processing of relevant information. For example, the automotive industry can use vehicle-mounted cameras for environmental perception, data fusion, and remote driving, while roadside cameras are used for road monitoring and traffic flow control. However, due to the limitations of current network communication technology, when transmitting multiple video data remotely or in real-time, problems such as channel overload, high transmission latency, and excessively long video encoding and decoding times arise. When the receiving entity does not have high requirements for the actual video data, the current mainstream solution is to convert all relevant target objects into animations or color blocks before transmission, reducing the impact of unnecessary information on the transmission rate.

[0003] However, if the original target in the video is completely converted into an animated or color block substitute, the difference in the marking effect between the original target and the substitute reduces the transmission of effective information. This may lead to an increase in driving risks for the recipient due to misidentification or omission of important information. Therefore, a more effective technical solution is needed. Summary of the Invention

[0004] To address the problems of existing technologies, this application provides a video image processing method, apparatus, device, and storage medium. The technical solution is as follows:

[0005] On one hand, a video image processing method is provided, the method comprising:

[0006] Acquire a first target frame video image, wherein the first target frame video image is any frame video image among the multiple frames of video images to be processed;

[0007] Target detection is performed on the first target frame video image to determine at least one target object in the first target frame video image;

[0008] Based on the preset classification rules for the target objects to be processed, the first target object to be processed among the at least one target objects is determined;

[0009] In the first target frame video image, the first target object to be processed is replaced with a preset target substitute to obtain the second target frame video image;

[0010] Wherein, the amount of data of the preset target substitute is less than the amount of data of the first target to be processed.

[0011] On the other hand, a video image processing apparatus is provided, the apparatus comprising:

[0012] The video image acquisition module is used to acquire a first target frame video image, wherein the first target frame video image is any one of the multiple frames of video images to be processed.

[0013] The target detection module is used to perform target detection on the first target frame video image and determine at least one target object in the first target frame video image;

[0014] The target classification module is used to determine the first target object to be processed among the at least one target objects based on a preset classification rule for target objects to be processed;

[0015] The target substitution module is used to replace the first target object to be processed with a preset target substitute in the first target frame video image to obtain a second target frame video image;

[0016] Wherein, the amount of data of the preset target substitute is less than the amount of data of the first target to be processed.

[0017] On the other hand, a video image processing device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the video image processing method as described above.

[0018] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the video image processing method as described above.

[0019] The video image processing method, apparatus, device, and storage medium provided in this application have the following technical effects:

[0020] The technical solution provided in this application performs target detection and classification on video images, retaining important targets among all targets while converting other targets into smaller substitutes. The two are then combined and output within a short time delay. On the one hand, this does not affect the actual output effect of the video, ensuring that information can be transmitted accurately and in a timely manner. On the other hand, it reduces the amount of video data, increases the video transmission rate, and reduces the video transmission latency. Attached Figure Description

[0021] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic flowchart of a video image processing method provided in an embodiment of this application;

[0023] Figure 2 This is a flowchart illustrating a method for determining a target object to be processed, provided in an embodiment of this application.

[0024] Figure 3 This is a schematic diagram of a process for replacing a target object to be processed with a preset target substitute, provided in an embodiment of this application;

[0025] Figure 4 This is a schematic diagram of another process for replacing the target object to be processed with a preset target substitute, provided in an embodiment of this application;

[0026] Figure 5 This is a schematic diagram of a process for tracking a second target object provided in an embodiment of this application;

[0027] Figure 6 This is a schematic diagram of a video image processing device provided in an embodiment of this application;

[0028] Figure 7 This is a schematic diagram of a target classification module device provided in an embodiment of this application;

[0029] Figure 8 This is a schematic diagram of a target replacement module provided in an embodiment of this application;

[0030] Figure 9 This is a schematic diagram of another target replacement module provided in an embodiment of this application;

[0031] Figure 10 This is a hardware structure block diagram of a video image processing server provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged in real time where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0034] The following describes a vehicle warning method provided by an embodiment of this application. Figure 1 This is a flowchart illustrating a vehicle warning method provided in an embodiment of this application. It should be noted that this specification provides the operational steps of the method as described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many steps and does not represent the only execution order. In actual system or product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 1 As shown, the above method may include:

[0035] S101, Obtain a first target frame video image, wherein the first target frame video image is any frame video image among the multiple frames of video images to be processed.

[0036] In the embodiments of this specification, specifically, the multi-frame video images to be processed can be multi-frame video images in video data. The video data can include real-time video data of the vehicle end acquired by the vehicle-mounted camera for visual perception or visual monitoring. The first target frame video image can be any frame of real-time video image in the multi-frame real-time video images of the current vehicle end.

[0037] S103, Target detection is performed on the first target frame video image to determine at least one target object in the first target frame video image.

[0038] In the embodiments of this specification, the above-mentioned target detection of the first target frame video image to determine at least one target object in the first target frame video image may include:

[0039] The first target frame video image is input into the target detection model for target detection to obtain a first target detection result, which includes at least one target object in the first target frame video image.

[0040] In a specific embodiment, the above-mentioned object detection model can be obtained by training a preset machine learning model on sample video images labeled with object tags. Specifically, the training method of the above-mentioned object detection model may include:

[0041] 1) Obtain video images of the sample vehicle labeled with the target object;

[0042] In practical applications, before performing neural network machine learning, training data can be determined first. Specifically, in the embodiments of this specification, sample video images labeled with target object tags can be obtained as training data.

[0043] Specifically, the aforementioned sample vehicle-mounted video images may include vehicle-mounted video images containing corresponding target objects. The aforementioned target object tags can serve as identifiers for the corresponding target objects. The aforementioned target objects can be objects related to the actual perception or monitoring needs of the vehicle-mounted video images; specifically, the aforementioned target objects may include, but are not limited to, roadside buildings, roadside equipment, pedestrians, and vehicles.

[0044] 2) Based on the above sample video images, a preset machine learning model is used for target detection training. During the target detection training, the model parameters of the preset machine learning model are adjusted until the target detection results output by the preset machine learning model match the above target labels.

[0045] Specifically, the preset machine learning model may include, but is not limited to, a neural network machine learning model. The model parameters may include the model parameters (weights) learned during training. The target detection results include the target objects in the sample video images.

[0046] 3) Use the machine learning model corresponding to the current model parameters as the target detection model mentioned above.

[0047] As can be seen from the above embodiments of this specification, this embodiment uses sample vehicle-side video images labeled with target object tags as training data. Through machine learning, the trained target detection model can detect target object tags in vehicle-side video images of the same type as the training data.

[0048] In the embodiments of this specification, the first target detection result may further include: the type information, first location information, and first physical attribute information of at least one target object.

[0049] Specifically, during the training process of the target detection model, the target label can also include the target type information, location information and physical attribute information. By training the target detection model with sample vehicle video images labeled with the target label, the target detection result of the target detection model can also include: the target type information, location information and physical attribute information.

[0050] Specifically, type information represents the basic classification category of the target object, and type information may include, but is not limited to: buildings, streetlights, traffic lights, trees, pedestrians, and vehicles. Location information represents the position of the target object relative to the current vehicle in the video image; the first location information can be the position information of the target object in the first target frame video image. Physical attribute information represents the physical attributes of the target object in the video image, and physical attribute information may include, but is not limited to, contour feature information; the first physical attribute information can be the physical attribute information of the target object in the first target frame video image.

[0051] S105, based on the preset classification rules for the target objects to be processed, determine the first target object to be processed among the above at least one target object.

[0052] In the embodiments of this specification, the first target object to be processed may be a target object in the first target frame video image that is unrelated to or weakly related to the current vehicle's driving path.

[0053] In a specific embodiment, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating a method for determining a first target object to be processed, as provided in an embodiment of this specification. Specifically, it may include:

[0054] S201, based on the first location information and type information of each target object, determine the first influence factor corresponding to each target object.

[0055] Specifically, the influence factor characterizes the degree to which the location and type information of a target object in the current video image affects the current vehicle's driving path. Generally, the larger the influence factor, the greater the influence. In the embodiments of this specification, an influence factor algorithm can be obtained by summarizing the location and type information of a large number of target objects and the corresponding preset influence factors. Based on the influence factor algorithm, the location and type information of the target object are analyzed to obtain the influence factor of the target object.

[0056] In the embodiments described in the specification, the first influence factor can be an influence factor in the first target frame video image.

[0057] S203, the target object whose first influence factor among the above at least one target object satisfies the first preset condition is taken as the first target object to be processed.

[0058] Specifically, in the embodiments of this specification, the influencing factor may include, but is not limited to, irrelevant, weakly correlated, and strongly correlated factors. Based on actual monitoring needs and vehicle safety warning needs, the influencing factor being irrelevant or weakly correlated is set as the aforementioned first preset condition. In the embodiments of this specification, the aforementioned first target object to be processed may be a target object whose first influencing factor is irrelevant or weakly correlated.

[0059] In practical applications, the first target to be processed can be a fixed target that is not related to the vehicle's planned path or actual driving path, or a static or dynamic target with little relevance. Specifically, the first target to be processed can include, but is not limited to, buildings, streetlights, traffic lights, curbs, pedestrians on the curb, and vehicles parked on the roadside.

[0060] S107, in the first target frame video image, the first target object to be processed is replaced with a preset target substitute to obtain a second target frame video image. The data size of the preset target substitute is less than the data size of the first target object to be processed.

[0061] Specifically, the aforementioned preset target substitute can be a pre-set target substitute that matches the type information and physical attribute information of the aforementioned first target to be processed, and the data volume of the aforementioned preset target substitute is less than the data volume of the aforementioned first target to be processed.

[0062] In an optional embodiment, such as Figure 3 As shown, replacing the first target object to be processed with a preset target substitute in the first target frame video image to obtain the second target frame video image may include:

[0063] S301, based on the first location information of the first target object to be processed, semantic segmentation is performed on the first target object to be processed in the first target frame video image to obtain the segmentation region corresponding to the first target object to be processed.

[0064] In practical applications, semantic segmentation involves classifying each pixel in a video image into a corresponding category, thus achieving pixel-level classification.

[0065] Specifically, based on the first location information of the first target object to be processed, semantic segmentation is performed on the first target object in the first target frame video image to determine the region where the original pixel image of the first target object is located, and the region where the original pixel image of the first target object is located is taken as the segmentation region corresponding to the first target object.

[0066] S303, based on the type information and first physical attribute information of the first target object to be processed, determine the preset target substitute corresponding to the first target object to be processed.

[0067] Specifically, a preset target substitute is determined that matches the type information and first physical attribute information of the first target object to be processed. That is, the type information and first physical attribute information of the first target object to be processed can be identified through the preset target substitute. The aforementioned preset target substitute may include, but is not limited to, animated cartoons or color blocks with a small amount of data.

[0068] S305, in the corresponding segmented region, the first target object to be processed is replaced with the corresponding preset target substitute to obtain the replaced first target frame video image.

[0069] Specifically, in the segmented region corresponding to the first target frame video image, the first target object to be processed is replaced with a preset animated cartoon or color block to obtain the replaced first target frame video image. The data size of the replaced first target frame video image is smaller than that of the first target frame video image.

[0070] S307, in the replaced first target frame video image, the edge contours of the corresponding segmented regions are smoothed to obtain the second target frame video image.

[0071] In practical applications, since the edge contours of the segmented region are relatively sharp, and the contours of the preset target substitute may not completely coincide with the edge contours of the segmented region, it is necessary to blur and smooth the edge contours to make the edge transitions more natural.

[0072] As can be seen from the above embodiments of this specification, this embodiment, while retaining the location information and physical attribute information of the first target object to be processed, replaces the first target object to be processed with a preset target substitute with a smaller data volume, thereby reducing the data volume of the video image without affecting the actual output effect.

[0073] In another alternative embodiment, such as Figure 4 As shown, when the first target object to be processed includes multiple first target objects to be processed, the step of replacing the first target object to be processed with a preset target substitute in the first target frame video image to obtain the second target frame video image may further include:

[0074] S401, based on the first position information of the plurality of first target objects to be processed, instance segmentation is performed on the plurality of first target objects to be processed in the first target frame video image to obtain a plurality of segmented regions corresponding to the plurality of first target objects to be processed.

[0075] In practical applications, instance segmentation not only performs pixel-level classification, but also distinguishes different instances based on specific categories, with each instance being a concrete object of a class.

[0076] Specifically, when the first position information of the above-mentioned multiple first target objects to be processed is used to perform instance segmentation of the multiple target objects to be processed in the first target frame video image, the original pixel image regions of the multiple first target objects to be processed are determined respectively, and the original pixel image regions of the multiple first target objects to be processed are used as the segmentation regions corresponding to the multiple first target objects to be processed.

[0077] S403, based on the type information and first physical attribute information of the aforementioned multiple first target objects to be processed, determine the multiple preset target substitutes corresponding to the aforementioned multiple first target objects to be processed.

[0078] Specifically, multiple preset target substitutes are determined to match the type information and first physical attribute information of multiple first target objects to be processed. That is, the type information and first physical attribute information of the corresponding multiple first target objects to be processed can be identified through multiple preset target substitutes. The aforementioned preset target substitutes may include, but are not limited to, animated cartoons or color blocks with a small amount of data.

[0079] In the embodiments of this specification, when the above-mentioned plurality of first target objects to be processed include a plurality of first target objects to be processed of the same type, the plurality of preset target substitutes corresponding to the plurality of first target objects to be processed of the same type are set as a plurality of animated cartoons or color blocks containing the same type of information but different style information.

[0080] In practical applications, the aforementioned style information may include, but is not limited to, color information and shadow information.

[0081] S405, in the corresponding multiple segmented regions, the multiple first target objects to be processed are replaced with the corresponding multiple preset target substitutes respectively, to obtain the replaced first target frame video image.

[0082] Specifically, in the segmented region corresponding to the first target frame video image, multiple first target objects to be processed are replaced with multiple corresponding animated cartoons or color blocks to obtain the replaced first target frame video image. The data size of the replaced first target frame video image is smaller than that of the first target frame video image.

[0083] S407, in the replaced first target frame video image, the edge contours of the corresponding multiple segmented regions are smoothed to obtain the second target frame video image.

[0084] Specifically, the smoothing of the edge contours of multiple segmented regions can be found in the relevant description of S407 above, and will not be repeated here.

[0085] As can be seen from the above embodiments of this specification, this embodiment replaces multiple first target objects to be processed with multiple corresponding preset target substitutes with smaller data volumes. While retaining the location information and physical attribute information of multiple first target objects to be processed, it distinguishes multiple first target objects to be processed that belong to the same type, thereby reducing the data volume of the video frame and reducing the transmission latency of the video frame.

[0086] In a specific embodiment, such as Figure 5 As shown, when the first preset condition includes the second preset condition, after selecting the target object whose first influence factor among the at least one target object satisfies the first preset condition as the first target object to be processed, the method may further include:

[0087] S501, the target objects whose first influencing factors satisfy the second preset conditions among the first target objects to be processed are taken as the second target objects to be processed.

[0088] Specifically, in the embodiments of this specification, based on actual monitoring needs and vehicle safety warning needs, the influencing factor is set as the second preset condition mentioned above as weakly correlated. The second target object to be processed can be a target object whose first influencing factor is weakly correlated.

[0089] In practical applications, the second target object to be processed can be a static or dynamic target object that has little correlation with the vehicle's planned path or actual driving path. Specifically, the second target object to be processed can include, but is not limited to, pedestrians on the curb and vehicles parked on the roadside.

[0090] Accordingly, after replacing the first target object to be processed with a preset target substitute in the first target frame video image to obtain the second target frame video image, the method may further include:

[0091] S503, Obtain the next frame video image of the first target frame video image mentioned above.

[0092] In practical applications, based on the timeline order of the multiple video frames to be processed, the next frame of the first target video frame is obtained.

[0093] S505, input the above-mentioned next frame video image into the above-mentioned target detection model for target detection, and obtain the second target detection result.

[0094] Specifically, the target detection for the next frame of video image is similar to the target detection for the first target frame of video image in S103. For details, please refer to the relevant description of the target detection for the first target frame of video image in S103, which will not be repeated here.

[0095] S507, when the second target detection result includes the second target to be processed, the second target detection result also includes the second location information of the second target to be processed.

[0096] Specifically, the second location information represents the location information of the second target object in the next frame of video image.

[0097] S509, based on the type information and second location information of the second target object to be processed, determine the second influence factor of the second target object to be processed.

[0098] Specifically, the second influence factor characterizes the influence of the second target object on the current vehicle's driving path in the next frame of the video image. Specifically, the steps for determining the second influence factor of the second target object are similar to those for determining the first influence factor of the target object in S201. For specific steps, please refer to the relevant description of determining the first influence factor of the target object in S201, which will not be repeated here.

[0099] S511, determine whether the second influencing factor satisfies the first preset condition.

[0100] Specifically, target tracking is performed on the second target object in the first target frame video image, and it is determined whether the second target object can still be identified as the target object in the next frame video image.

[0101] S513, when the determination is negative, the preset target substitute corresponding to the second target to be processed is replaced with the second target object to be processed.

[0102] In practical applications, as the movement path of the second target object changes, the influence factor of the second target object also changes. When the second influence factor of the second target object does not meet the first preset condition, the current second target object has a greater influence on the current vehicle's driving path. Therefore, instead of replacing the current second target object with the corresponding preset target substitute, the real-time original image of the current second target object is directly transmitted to ensure that important information in the video image can be transmitted accurately and in a timely manner.

[0103] This application provides a video image processing apparatus, such as... Figure 6 As shown, the above-mentioned device includes:

[0104] The video image acquisition module 610 is used to acquire a first target frame video image, wherein the first target frame video image is any one of the multiple frames of video images to be processed.

[0105] The target detection module 620 is used to perform target detection on the first target frame video image and determine at least one target object in the first target frame video image;

[0106] The target classification module 630 is used to determine the first target object to be processed among the above-mentioned at least one target object based on a preset target object classification rule;

[0107] The target replacement module 640 is used to replace the first target object to be processed with a preset target substitute in the first target frame video image to obtain the second target frame video image.

[0108] The amount of data for the aforementioned preset target substitute is less than the amount of data for the aforementioned first target to be processed.

[0109] In the embodiments described in this specification, the target detection module 620 may include:

[0110] The first target detection result unit is used to input the first target frame video image into the target detection model for target detection and obtain a first target detection result, wherein the first target detection result includes at least one target object in the first target frame video image.

[0111] In the embodiments of this specification, the first target detection result may further include: type information and first location information of each of the at least one target.

[0112] In a specific embodiment, such as Figure 7 As shown, the target classification module 630 described above may include:

[0113] The first impact factor unit 631 is used to determine the first impact factor corresponding to each of the above target objects based on the first location information and type information of each target object.

[0114] The first target object unit 632 is used to select the target object whose corresponding first influence factor meets the first preset condition from the at least one target object as the first target object to be processed.

[0115] In an optional embodiment, such as Figure 8 As shown, the first target detection result may also include the first physical attribute information of the first target object to be processed.

[0116] The aforementioned target replacement module 640 may include:

[0117] The first target segmentation unit 641 is used to perform semantic segmentation on the first target object in the first target frame video image based on the first location information of the first target object to be processed, so as to obtain the segmentation region corresponding to the first target object to be processed.

[0118] The first preset target substitute determination unit 642 is used to determine the preset target substitute corresponding to the first target to be processed based on the type information and the first physical attribute information of the first target to be processed.

[0119] The first preset target substitute replacement unit 643 is used to replace the first target object to be processed with the corresponding preset target substitute in the corresponding segmented area to obtain the replaced first target frame video image.

[0120] The first edge contour processing unit 644 is used to smooth the edge contour of the corresponding segmented region in the replaced first target frame video image to obtain the second target frame video image.

[0121] In another alternative embodiment, such as Figure 9 As shown, when the first target to be processed includes multiple first target to be processed objects, the first target detection result may also include the first physical attribute information of the multiple first target to be processed objects.

[0122] The aforementioned target replacement module 640 may further include:

[0123] The second target segmentation unit 645 is used to perform instance segmentation of the multiple first target objects in the first target frame video image based on the first location information of the multiple first target objects to be processed, so as to obtain multiple segmentation regions corresponding to the multiple first target objects to be processed.

[0124] The second preset target substitute determination unit 646 is used to determine multiple preset target substitutes corresponding to the multiple first target objects to be processed based on the type information and first physical attribute information of the multiple first target objects to be processed.

[0125] The second preset target substitute replacement unit 647 is used to replace the plurality of first target objects to be processed with the plurality of preset target substitutes in the corresponding plurality of segmented regions to obtain the replaced first target frame video image.

[0126] The second edge contour processing unit 648 is used to smooth the edge contours of the corresponding multiple segmented regions in the replaced first target frame video image to obtain the second target frame video image.

[0127] In a specific embodiment, when the first preset condition includes the second preset condition, the device may further include:

[0128] The second target object unit is used to take the target objects whose corresponding first influence factors meet the second preset conditions from the first target objects to be processed as the second target objects to be processed.

[0129] The next frame video image acquisition unit is used to acquire the next frame video image of the first target frame video image mentioned above;

[0130] The second target detection result unit is used to input the above-mentioned next frame video image into the above-mentioned target detection model to perform target detection and obtain the second target detection result;

[0131] The second location information unit is configured to, when the second target detection result includes the second target object to be processed, also include the second location information of the second target object to be processed in the second target detection result.

[0132] The second impact factor unit is used to determine the second impact factor of the second target object based on the type information and the second location information of the second target object to be processed.

[0133] The first preset condition judgment unit is used to determine whether the second influence factor satisfies the first preset condition.

[0134] The second target object replacement unit is used to replace the preset target substitute corresponding to the second target object with the second target object when the determination is negative.

[0135] The apparatus and method embodiments described above are based on the same inventive concept.

[0136] This application provides a video image processing device, which includes a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the video image processing method provided in the above method embodiments.

[0137] Memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for the functions, etc.; the data storage area can store data created based on the use of the aforementioned devices. Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.

[0138] The methods and embodiments provided in this application can be executed in a mobile terminal, computer terminal, server, or similar computing device; that is, the aforementioned computer device may include a mobile terminal, computer terminal, server, or similar computing device. Taking running on a server as an example... Figure 10 This is a hardware structure block diagram of a video image processing server provided in an embodiment of this application. For example... Figure 10 As shown, the video image processing server 1000 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1010 (CPUs 1010 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 1030 for storing data, and one or more storage media 1020 (e.g., one or more mass storage devices) for storing application programs 1023 or data 1022. The memory 1030 and storage media 1020 may be temporary or persistent storage. The program stored in the storage media 1020 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 1010 may be configured to communicate with the storage media 1020 and execute the series of instruction operations in the storage media 1020 on the video image processing server 1000. The video image processing server 1000 may also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1040, and / or one or more operating systems 1021, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0139] The input / output interface 1040 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the video image processing server 1000. In one example, the input / output interface 1040 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 1040 can be a radio frequency (RF) module for wireless communication with the Internet.

[0140] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, the video image processing server 1000 may also include... Figure 10The more or fewer components shown, or having the same Figure 10 The different configurations shown.

[0141] This application embodiment also provides a storage medium, which can be disposed in a server to store at least one instruction or at least one program related to implementing a video image processing method of one of the method embodiments. The at least one instruction or the at least one program is loaded and executed by the processor to implement the video image processing method provided by the above method embodiment.

[0142] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0143] As can be seen from the embodiments of the video image processing method, apparatus, device, or storage medium provided in this application, the technical solution provided in this application performs target detection and classification on the video image, retains important targets among all targets, and transforms other targets to be processed into substitutes with smaller data volumes. The two are combined and output within a short time delay. On the one hand, this does not affect the actual output effect of the video, ensuring that important information can be transmitted accurately and in a timely manner. On the other hand, it reduces the data volume of the video, increases the video transmission rate, and reduces the video transmission delay. It can also perform target tracking on weakly correlated targets among other targets to be processed. When a weakly correlated target is transformed into a strongly correlated target, the real-time original image of the weakly correlated target is directly output, further ensuring the accurate transmission of important information in the video image.

[0144] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0145] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0146] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing the relevant hardware to implement them. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0147] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A video image processing method, characterized in that, The method includes: Acquire a first target frame video image, wherein the first target frame video image is any frame video image among the multiple frames of video images to be processed; Target detection is performed on the first target frame video image to determine at least one target object in the first target frame video image; Based on a preset classification rule for the target objects to be processed, a first target object to be processed is determined among the at least one target objects; the first target object to be processed is a target object in the first target frame video image that is unrelated to or weakly related to the current vehicle's driving path; In the first target frame video image, the first target object to be processed is replaced with a preset target substitute to obtain a second target frame video image; the preset target substitute is a target substitute that is pre-set to match the type information and physical attribute information of the first target object to be processed; the physical attribute information includes contour feature information. Wherein, the amount of data of the preset target substitute is less than the amount of data of the first target to be processed; The step of performing target detection on the first target frame video image to determine at least one target object in the first target frame video image includes: The first target frame video image is input into the target detection model for target detection to obtain a first target detection result, which includes at least one target object in the first target frame video image. The first target detection result also includes type information and first location information for each of the at least one target object; The determination of the first target object among the at least one target objects based on the preset target object classification rules includes: Based on the first location information and type information of each target object, determine the first influence factor corresponding to each target object; The target object whose first influence factor among the at least one target object meets the first preset condition is designated as the first target object to be processed.

2. The method according to claim 1, characterized in that, The first target detection result also includes the first physical attribute information of the first target object to be processed; The step of replacing the first target object to be processed with a preset target substitute in the first target frame video image to obtain the second target frame video image includes: Based on the first location information of the first target object to be processed, semantic segmentation is performed on the first target object in the first target frame video image to obtain the segmentation region corresponding to the first target object to be processed. Based on the type information and the first physical attribute information of the first target object to be processed, a preset target substitute corresponding to the first target object to be processed is determined; In the corresponding segmented region, the first target object to be processed is replaced with the corresponding preset target substitute to obtain the replaced first target frame video image; In the replaced first target frame video image, the edge contours of the corresponding segmented regions are smoothed to obtain the second target frame video image.

3. The method according to claim 1, characterized in that, When the first target object to be processed includes multiple first target objects to be processed, the first target detection result also includes the first physical attribute information of the multiple first target objects to be processed; The step of replacing the first target object to be processed with the corresponding preset target substitute in the corresponding segmented region to obtain the second target frame video image includes: Based on the first location information of the plurality of first target objects to be processed, instance segmentation is performed on the plurality of first target objects to be processed in the first target frame video image to obtain a plurality of segmented regions corresponding to the plurality of first target objects to be processed. Based on the type information and first physical attribute information of the plurality of first target objects to be processed, a plurality of preset target substitutes corresponding to the plurality of first target objects to be processed are determined respectively; In the corresponding multiple segmentation regions, the multiple first target objects to be processed are respectively replaced with the corresponding multiple preset target substitutes to obtain the replaced first target frame video image; In the replaced first target frame video image, the edge contours of the corresponding multiple segmented regions are smoothed to obtain the second target frame video image.

4. A video image processing apparatus, characterized in that, The device includes: The video image acquisition module is used to acquire a first target frame video image, wherein the first target frame video image is any one of the multiple frames of video images to be processed. The target detection module is used to perform target detection on the first target frame video image and determine at least one target object in the first target frame video image; The target classification module is used to determine a first target object among the at least one target objects based on a preset target object classification rule; the first target object is a target object in the first target frame video image that is unrelated to or weakly related to the current vehicle's driving path; The target replacement module is used to replace the first target object to be processed with a preset target substitute in the first target frame video image to obtain a second target frame video image; the preset target substitute is a target substitute that is pre-set to match the type information and physical attribute information of the first target object to be processed; the physical attribute information includes contour feature information. Wherein, the amount of data of the preset target substitute is less than the amount of data of the first target to be processed; The target detection module includes: The first target detection result unit is used to input the first target frame video image into the target detection model for target detection and obtain a first target detection result, wherein the first target detection result includes at least one target object in the first target frame video image. The first target detection result further includes: type information and first location information of each of the at least one target object; The target classification module includes: The first impact factor unit is used to determine the first impact factor corresponding to each target object based on the first location information and type information of each target object; The first target object unit is used to select the target object whose corresponding first influence factor meets the first preset condition from the at least one target object as the first target object to be processed.

5. The apparatus according to claim 4, characterized in that, The first target detection result also includes the first physical attribute information of the first target object to be processed; The target replacement module includes: The first target segmentation unit is used to perform semantic segmentation on the first target object in the first target frame video image based on the first location information of the first target object to be processed, so as to obtain the segmentation region corresponding to the first target object to be processed. The first preset target substitute determination unit is used to determine the preset target substitute corresponding to the first target object based on the type information and the first physical attribute information of the first target object to be processed; The first preset target substitute replacement unit is used to replace the first target object to be processed with the corresponding preset target substitute in the corresponding segmentation area to obtain the replaced first target frame video image; The first edge contour processing unit is used to smooth the edge contour of the corresponding segmented region in the replaced first target frame video image to obtain the second target frame video image.

6. The apparatus according to claim 4, characterized in that, When the first target object to be processed includes multiple first target objects to be processed, the first target detection result also includes the first physical attribute information of the multiple first target objects to be processed; The target replacement module also includes: The second target segmentation unit is used to perform instance segmentation on the plurality of first target objects to be processed in the first target frame video image based on the first location information of the plurality of first target objects to be processed, so as to obtain a plurality of segmentation regions corresponding to the plurality of first target objects to be processed. The second preset target substitute determination unit is used to determine multiple preset target substitutes corresponding to the multiple first target objects to be processed based on the type information and first physical attribute information of the multiple first target objects to be processed; The second preset target substitute replacement unit is used to replace the plurality of first target objects to be processed with the plurality of corresponding preset target substitutes in the corresponding plurality of segmented regions, so as to obtain the replaced first target frame video image; The second edge contour processing unit is used to smooth the edge contours of the corresponding multiple segmented regions in the replaced first target frame video image to obtain the second target frame video image.

7. A video image processing device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the video image processing method as described in any one of claims 1 to 3.

8. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the video image processing method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Video processing method and device, readable storage medium and electronic equipment

    CN112541870A