Scenery video generation method and device, computer equipment and storage medium

By identifying key features and category score predictions in the environmental image sequence of the on-board camera, high-quality landscape videos are generated, which solves the problem that existing on-board camera systems are difficult to automatically identify and shoot scenery, and improves the driving experience.

CN120302104APending Publication Date: 2025-07-11ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510341596.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing on-board camera system lacks the ability to automatically identify scenery along the way and advanced visual algorithms, making it difficult to accurately identify and automatically shoot and record landmark landscapes such as characteristic buildings, limiting the improvement of the driving experience.

Method used

By obtaining environmental image sequences, identifying key image features and calculating category scores, predicting target landscape categories, generating high-quality landscape videos, extracting maximum value features using sliding windows, combining weight matrix and association relationships, filtering and generating landscape videos.

Benefits of technology

It realizes accurate identification and automatic shooting of scenery along the way, generates high-quality landscape videos, improves driving experience and visual memory, and solves the limitations of the existing vehicle-mounted camera system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302104A_ABST
    Figure CN120302104A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicles, and discloses a landscape video generation method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining an environment image sequence collected in the driving process of a target vehicle; identifying a key image feature of each environment image in the environment image sequence, and obtaining a category score of a scenery sub-category corresponding to the key image feature; predicting a target scenery category hit by the environment image based on an association relationship between a preset scenery sub-category and a scenery category and the category score corresponding to the scenery sub-category; and obtaining a target environment image associated with the target scenery category from the environment image sequence, and generating a scenery video corresponding to the target scenery category by using the target environment image. According to the invention, the problem that the scenery along the way cannot be accurately identified and automatically shot due to the lack of automatic identification capability and advanced vision algorithm in the automatic scenery shooting function of the existing vehicle-mounted camera is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicles, and in particular to a method, device, computer device and storage medium for generating scenic videos. Background Art

[0002] With the development and popularization of intelligent vehicle technology, in-vehicle cameras have become one of the common and indispensable configurations in modern vehicles. In actual application scenarios, in-vehicle cameras carry multiple functions. On the one hand, they are committed to improving driving safety. For example, as a driving recorder, they can record the road conditions during driving in real time and provide evidence for possible traffic accidents; or as part of the assisted driving system, they can assist the driver in perceiving the surrounding environment and making accurate driving decisions in a timely manner. On the other hand, in-vehicle cameras are also constantly expanding their application boundaries and are gradually being used to enhance the driving experience. For example, by recording various landscapes and interesting events during the journey, they can add fun to the driver's journey.

[0003] However, there are still several limitations in the existing in-vehicle camera system in terms of the automatic landscape shooting function, which are specifically manifested as follows: The existing in-vehicle cameras lack the ability to automatically identify the landscapes along the way and trigger the shooting action, resulting in the need for users to manually operate to capture the landscapes of interest. Secondly, although some in-vehicle cameras incorporate automation and intelligence elements, they lack more advanced vision algorithms, making it difficult to achieve more accurate landscape recognition. For example, for different types of characteristic buildings, such as European-style buildings, magnificent bridges, towering skyscrapers, and Ferris wheels and other iconic landscapes, it is difficult to accurately identify and automatically shoot and record them, which further limits the further expansion and optimization of the in-vehicle camera system in enhancing the driving experience. Summary of the Invention

[0004] In view of this, the embodiments of the present invention provide a method, device, computer device and storage medium for generating scenic videos to solve the problem that the existing in-vehicle cameras lack automatic recognition ability and advanced vision algorithms in the automatic landscape shooting function, resulting in the inability to accurately identify and automatically shoot the landscapes along the way.

[0005] In a first aspect, the embodiments of the present invention provide a method for generating a scenic video, the method comprising:

[0006] Obtaining an environmental image sequence collected during the driving of a target vehicle;

[0007] Identifying key image features of each environmental image in the environmental image sequence, and obtaining category scores of the landscape subcategories corresponding to the key image features;

[0008] Predict the target scenic category hit by the environmental image based on the association relationship between the preset scenic sub-categories and the scenic categories and the category scores corresponding to the scenic sub-categories;

[0009] Obtain the target environmental image associated with the target scenic category from the environmental image sequence, and generate a scenic video corresponding to the target scenic category using the target environmental image.

[0010] Further, the identification of the key image features in the environmental image includes:

[0011] Identify the original image features in the environmental image;

[0012] Use a sliding window to extract the element maximum value in the original image features until the spatial dimension of the original image features in the environmental image reaches the specified dimension, and take the element maximum value as the key image feature.

[0013] Further, the obtaining of the category scores of the scenic sub-categories corresponding to the key image features includes:

[0014] Obtain the weight matrix corresponding to the scenic sub-category;

[0015] Calculate the category scores of the scenic sub-categories corresponding to the key image features according to the key image features and the weight matrix.

[0016] Further, the calculation of the category scores of the scenic sub-categories corresponding to the key image features according to the key image features and the weight matrix includes:

[0017] Multiply the key image features by the weight matrix to obtain the initial category scores corresponding to the scenic sub-categories;

[0018] Analyze the spatial distribution of the key image feature vectors in the environmental image to obtain an analysis result;

[0019] If the analysis result is that the key image feature vectors are concentrated, correct the initial category scores of the scenic sub-categories according to the first correction strategy to obtain the category scores; or, if the analysis result is that the key image feature vectors are dispersed, correct the initial category scores of the scenic sub-categories according to the second correction strategy to obtain the category scores, where the first correction strategy is used to perform a correction operation on the initial category scores according to the score improvement coefficient corresponding to the concentrated distribution, and the second correction strategy is used to perform a correction operation on the initial category scores according to the score reduction coefficient corresponding to the dispersed distribution.

[0020] Further, predicting the target scenic category hit by the environmental image based on the association relationship between the preset scenic sub - categories and the scenic categories, and the category scores corresponding to the scenic sub - categories includes:

[0021] Obtain the association relationship between the preset scenic sub - categories and the scenic categories;

[0022] Determine the scenic category associated with each scenic sub - category according to the association relationship;

[0023] Calculate the probability prediction value of each scenic category corresponding to the environmental image based on the category scores;

[0024] Take the scenic category corresponding to the maximum value among the probability prediction values as the target scenic category.

[0025] Further, calculating the probability prediction value of each scenic category corresponding to the environmental image based on the category scores includes:

[0026] Calculate the initial probability prediction value of each scenic category corresponding to the environmental image based on the category scores;

[0027] Detect whether there is an association relationship between each scenic sub - category in the scenic category;

[0028] If there is an association relationship, correct the initial probability prediction value of the environmental perception category according to the third correction strategy to obtain the probability prediction value, where the third correction strategy is used to perform a correction operation on the initial probability prediction value according to the tightness of the association relationship and the importance of the scenic sub - category in the environmental image.

[0029] Further, generating a scenic video corresponding to the target scenic category using the target environmental image includes:

[0030] Detect the image quality of the target environmental image to obtain the quality score of the target environmental image;

[0031] Filter the target environmental image based on the quality score to obtain candidate environmental images;

[0032] Match the candidate environmental images with corresponding video templates, and generate corresponding scenic videos based on the video templates and the candidate environmental images.

[0033] In a second aspect, an embodiment of the present invention provides a device for generating a scenic video, the device includes:

[0034] An acquisition module, configured to acquire a sequence of environmental images collected during the driving of a target vehicle;

[0035] An identification module, configured to identify key image features of each environmental image in the environmental image sequence and obtain category scores of the landscape subcategories corresponding to the key image features;

[0036] A prediction module, configured to predict the target landscape category hit by the environmental image based on the association relationship between preset landscape subcategories and landscape categories and the category scores corresponding to the landscape subcategories;

[0037] A generation module, configured to obtain the target environmental image associated with the target landscape category from the environmental image sequence and generate a landscape video corresponding to the target landscape category by using the target environmental image.

[0038] In a third aspect, an embodiment of the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any corresponding implementation manner thereof.

[0039] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the method according to the first aspect or any corresponding implementation manner thereof.

[0040] The method provided by the embodiments of the present application has the following beneficial effects:

[0041] The method provided by the embodiments of this application can provide a data basis for subsequent mining of scenic content and generation of high-quality scenic videos by obtaining the environmental image sequence collected during the driving of the target vehicle. Identifying key image features helps to extract representative and distinguishable elements from the complex environmental image sequence. Obtaining the category scores of the scenic subcategories corresponding to the key image features can quantify the degree of compliance of the key features. This quantification method can intuitively compare the possibilities of different scenic subcategories according to the score levels, providing a data basis for judging which scenic subcategory the environmental image is more inclined to. By presetting the association relationship between the scenic subcategories and the scenic categories and the category scores corresponding to the scenic subcategories to predict the target scenic category hit by the environmental image, on the one hand, the classification of the scenery is not limited to a single level, but can be reasonably summarized from the microscopic subcategories to the macroscopic scenic categories; on the other hand, it improves the accuracy of predicting the target scenic category, making the scenic category more in line with the actual scenery content of the image, and solving the problem that it is difficult for existing on-vehicle cameras to accurately identify the scenic categories. Accurately obtaining the target environmental images associated with the target scenic category from the environmental image sequence avoids the mixing of irrelevant images into the scenic video, ensures that the scenic video unfolds around a specific scenic category, and improves the professionalism and appreciation of the video. Generating a scenic video corresponding to the target scenic category using the target environmental images realizes the transformation from scattered image data to a complete and coherent video, makes up for the deficiency that it is difficult for existing on-vehicle camera systems to automatically shoot and generate high-quality scenic videos, improves the user experience during driving, and leaves a better visual memory for the journey. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 is a flowchart of the method for generating a scenic video according to an embodiment of the present invention;

[0044] Figure 2 is a schematic diagram of the association relationship between the scenic subcategories and the scenic categories according to an embodiment of the present invention;

[0045] Figure 3 is a structural block diagram of the visual model according to an embodiment of the present invention;

[0046] Figure 4 is a schematic diagram of the application process of the visual model according to an embodiment of the present invention;

[0047] Figure 5A flowchart of a method for configuring a vision model according to an embodiment of the present invention;

[0048] Figure 6 A structural block diagram of a generating device for a scenic video according to an embodiment of the present invention;

[0049] Figure 7 A schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] According to an embodiment of the present invention, there are provided a method, a device, a computer device, and a storage medium for generating a scenic video. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0052] In this embodiment, a method for generating a scenic video is provided. Figure 1 It is a flowchart of a method for generating a scenic video according to an embodiment of the present invention. As Figure 1 shown, the process includes the following steps:

[0053] Step S11, obtaining an environmental image sequence collected during the driving of a target vehicle.

[0054] In the embodiment of the present application, during the driving of the vehicle, an environmental image sequence can be obtained through relevant image acquisition devices equipped on the vehicle, such as a driving recorder (DVR). The driving recorder continuously takes images at a specific time interval or specific trigger condition during the driving of the vehicle, and these captured multiple images constitute the environmental image sequence. For example, the image acquisition device can take a picture every 1 second, or perform frame extraction on the video within 5 minutes of the vehicle's driving to convert it into pictures, etc., to collect rich image information corresponding to different times and positions along the vehicle's route, thereby forming a continuous image sequence reflecting the vehicle's surrounding environment situation, providing the most original data source for subsequent operations such as recognition, classification, and generation of scenic videos of the scenery along the way.

[0055] Step S12: Identify the key image features of each environmental image in the environmental image sequence, and obtain the category scores of the landscape subcategories corresponding to the key image features.

[0056] In the embodiment of the present application, identifying the key image features of each environmental image in the environmental image sequence includes the following steps A1 - A2:

[0057] Step A1: Identify the original image features in the environmental image.

[0058] Specifically, use a variety of image processing algorithms and techniques to mine the original image features in the environmental image. For example, use edge detection algorithms such as the Sobel operator and Canny edge detection. By calculating the change gradient of the pixel gray values in the image, identify the contour edges of the objects in the image. These edges can reflect the shape and structural features of the objects and are one of the important original image features. At the same time, use texture analysis algorithms such as the gray - level co - occurrence matrix (GLCM) method. By statistically analyzing the gray - level combinations of pixel pairs in different directions and at different distances in the image, obtain the texture features of the image, such as the roughness and directionality of the texture, which helps to distinguish objects with different materials or surface features, such as identifying a smooth lake surface or a rough rock surface. In addition, by calculating the color histograms of the image in different color spaces (such as RGB, HSV, etc.), obtain the color distribution information, such as the proportion of blue pixels in the image and the distribution range of different hues. This is very important for identifying landscape categories closely related to colors, such as blue sky and sunset glow. Through different algorithm operations, identify the original image features in the environmental image, providing a rich data basis for subsequent screening of key image features.

[0059] Step A2: Use a sliding window to extract the element maximum value from the original image features until the spatial dimension of the original image features in the environmental image reaches the specified dimension, and take the element maximum value as the key image feature.

[0060] Specifically, first, determine the size and shape of the sliding window. For example, it can be a 3x3 or 5x5 rectangular window. Then, starting from the upper left corner of the image, let the sliding window slide on the original image feature data according to the set step size (such as moving 1 pixel to the right or down each time). At each position of the sliding window, extract the maximum value of the original image feature elements within the window. For example, for the edge intensity feature, if there are multiple edge intensity values within a certain 3x3 window, select the maximum edge intensity value as the representative value at that window position. As the sliding window continuously traverses the entire image, a series of maximum values are gradually obtained. Continue this process while monitoring the change in the spatial dimension of the original image features. The spatial dimension can refer to the storage space size occupied by the feature data or forms such as the number of rows and columns of the feature matrix. When the spatial dimension reaches the preset specified dimension, stop the sliding window operation. The maximum values obtained at this time are determined as the key image features. For example, if the set specified dimension is to compress the spatial dimension of the original image features to 1 / 4 of the original, then through continuous sliding window extraction of maximum values, the amount of key image feature data finally obtained will be significantly reduced, and the most prominent and representative parts of the original image features are retained, enabling subsequent processing to focus on these key information, improving processing efficiency and accuracy, and at the same time reducing the burden of data storage and calculation.

[0061] It should be noted that the key image features refer to the significant features retained after spatial dimension compression, and these features simultaneously satisfy: (1) having a local maximum response value in spatial distribution; (2) maintaining class discrimination during the dimensionality reduction process; (3) conforming to the focus of human visual attention mechanism. For example, for a desert scene, the continuous maximum values extracted by the sliding window on the texture feature map form the trend feature of sand dune ripples; for a seaside scene, the maximum value distribution of the color feature presents the gradual transition feature of the sea water and the beach. This feature extraction method not only retains the spatial structure information of the original image but also strengthens the local features with class discriminability through max pooling.

[0062] Taking natural landscapes as an example, the key image features can refer to the following in specific embodiments:

[0063] Mountain edge features: When processing mountain landscape images, the sliding window slides on the edge feature map extracted by the convolutional neural network to capture the significant edges of the rock contours. For example, within a 3x3 window, the maximum edge response value corresponds to the sharp turn of the ridge line, and this high-gradient change area is extracted as the key terrain feature.

[0064] Water body color feature: For lake scenes, on the feature map in the HSV color space, the maximum hue value of a specific area may be captured by a 5x5 window. For example, the pixel group corresponding to the highest saturation value within a certain window characterizes the specular highlight feature of the lake surface reflection area.

[0065] Vegetation texture feature: In forest image processing, on the texture feature map extracted by the Gabor filter, the maximum texture response extracted by the sliding window corresponds to the dense texture area of the leaf clusters. For example, the maximum value of the histogram of oriented gradients (HOG) within a certain window indicates the arrangement feature of branches and leaves at a specific angle.

[0066] Building structure feature: When the image contains human landscapes, on the shape feature map, the window maximum value may capture the right-angle feature of the building outline. For example, the maximum response value of line detection within a certain window corresponds to the feature of the ancient building eaves.

[0067] The method provided by the embodiments of the present application for identifying the features of the original image can comprehensively capture various information bases in the image. By using a sliding window to extract the maximum value of elements in the features of the original image and determining the key image features when reaching the specified dimension, it helps to focus on the significant feature parts in the image, reduces the subsequent calculation amount while highlighting the feature information that has an important impact on landscape classification, enabling the model to efficiently and accurately judge and determine the scores of landscape subcategories based on the key features, and improving the accuracy and efficiency of overall landscape recognition.

[0068] In the embodiments of the present application, obtaining the category scores of the landscape subcategories corresponding to the key image features includes the following steps B1 - B2:

[0069] Step B1, obtaining the weight matrix corresponding to the landscape subcategory.

[0070] Specifically, the acquisition of the weight matrix is based on a large amount of previous data training and analysis. In the training stage, an image sample set containing numerous labeled landscape subcategories will be collected. These samples include various different landscape subcategory situations such as European-style architecture, lakeside roads, blue skies, etc., and the true landscape subcategory attribution corresponding to each image is clarified. Then, an initial classification model structure is built using machine learning algorithms or deep learning frameworks (such as TensorFlow, PyTorch, etc.). This structure can be a neural network model architecture or other model forms suitable for image classification tasks. The labeled image samples are input into this model for training. During the training process, the model will continuously adjust the parameters inside the model according to the features of the input images and the corresponding landscape subcategory labels, with the aim of enabling the model to determine the landscape subcategory to which the image belongs based on the image features. After multiple rounds of training iterations, when the model reaches good performance indicators such as classification accuracy on the validation set or test set, the set of weight parameters corresponding to each landscape subcategory in the model at this time is extracted and constructed into a weight matrix corresponding to the landscape subcategory in a specific data format (such as a two-dimensional matrix form, etc.).

[0071] As an example, for the landscape subcategory of European-style architecture, assuming that after model training, the weight value corresponding to the key image feature related to the building outline in its weight matrix is relatively high, indicating that when analyzing and judging in combination with the key image features of the actual image later, the building outline feature will play a relatively important role in determining whether the image belongs to the landscape subcategory of European-style architecture; while for relatively unimportant image features, their corresponding weight values will be lower.

[0072] Step B2, calculate the category score of the key image feature corresponding to the landscape subcategory according to the key image feature and the weight matrix.

[0073] The method provided by the embodiment of the present application can establish a quantitative correlation relationship between the key image feature and the landscape subcategory through the pre-set weight matrix, enabling the model to accurately evaluate the category score based on the importance differences of different features in different landscape subcategories, thereby providing reliable basic data for subsequent accurate prediction of the landscape category and enhancing the discrimination ability of the model for different landscape subcategories.

[0074] In the embodiment of the present application, step B2 includes the following steps B21 - B23:

[0075] Step B21, multiply the key image feature by the weight matrix to obtain the initial category score corresponding to the landscape subcategory.

[0076] Specifically, first, the key image features can be represented in the form of vectors. For example, assume that the key image feature vector of an image after extraction is x = [x1, x2,..., x n , where n represents the number of dimensions of the features, and each x i corresponds to the key image features in different aspects. The weight matrix W corresponding to the landscape subcategory is a two-dimensional matrix with dimensions m×n (where m represents the number of landscape subcategories and n is consistent with the key image feature dimension). Each element w ij in the matrix represents the weight value indicating the importance degree of the j-th key image feature for the i-th landscape subcategory. Then, according to the rules of matrix multiplication, multiply the key image feature vector x by the weight matrix W, that is, calculate z = x·W T (W T represents the transpose of the weight matrix), and the resulting z = [z1, z2,..., z m is the initial category score corresponding to each landscape subcategory. Among them, z i represents the initial score of the image for the i-th landscape subcategory.

[0077] Step B22: Analyze the spatial distribution of the key image feature vectors in the environmental image to obtain the analysis result.

[0078] Specifically, after obtaining the key image feature vectors, their distribution in the image space can be analyzed through various methods, including but not limited to: calculating distance metric indicators such as Euclidean distance and cosine similarity, and judging the degree of dispersion or aggregation according to the distance or similarity between vectors. A small distance means concentration, and a large distance difference means dispersion. K-Means and other clustering algorithms can also be used. Regarding the vectors as data points, judge according to the number of clusters and the density of vectors within the clusters. Fewer and denser clusters indicate a concentrated distribution, while more and sparser clusters indicate a dispersed distribution. From the perspective of image vision, the distribution of the positions of the objects corresponding to the key image features can also be analyzed. If they are concentrated in a local area, the vectors are probably concentrated, and if they are scattered at different positions, it is a dispersed distribution. Finally, the corresponding analysis result is obtained to provide a basis for score correction.

[0079] Step B23: If the analysis result shows that the key image feature vectors are concentrated, correct the initial category scores of the landscape subcategories according to the first correction strategy to obtain the category scores; or, if the analysis result shows that the key image feature vectors are dispersed, correct the initial category scores of the landscape subcategories according to the second correction strategy to obtain the category scores, where the first correction strategy is used to perform a correction operation on the initial category scores based on the score increase coefficient corresponding to the concentrated distribution, and the second correction strategy is used to perform a correction operation on the initial category scores based on the score decrease coefficient corresponding to the dispersed distribution.

[0080] Specifically, if the analysis result shows that the key image feature vectors are concentratedly distributed, indicating that the manifestation of the scenic subcategory is relatively typical, the first correction strategy at this time can be to increase the initial category score. A score increase coefficient σ (σ > 1, for example, taking a value of 1.2, etc.) can be set, and the initial category score z i is multiplied by this coefficient to obtain the corrected category score z i ′ = σ·z i , so that it has a relative advantage in subsequent judgments such as comparing with other scenic subcategory scores, and increases the probability of being judged as belonging to this scenic subcategory.

[0081] If the analysis result shows that the key image feature vectors are dispersedly distributed, indicating that the manifestation of the scenic subcategory is not typical, the second correction strategy at this time can be to decrease the initial category score. A score decrease coefficient γ (0 < γ < 1, for example, taking a value of 0.8, etc.) can be set, and the initial category score z i is multiplied by this coefficient to obtain the corrected category score z i ′ = γ·z i , so that it has a relative disadvantage in subsequent judgments such as comparing with other scenic subcategory scores, and decreases the probability of being misjudged as belonging to this scenic subcategory. In addition, it can be adjusted in combination with the specific situation of the dispersion. For example, if the dispersion degree exceeds the threshold, the score can be corrected according to a more stringent decrease rule. For example, in addition to multiplying by γ, a fixed offset δ (for example, δ = 0.3, etc.) is subtracted, that is, z i ′ = γ·z i - δ.

[0082] The method provided in the embodiment of the present application multiplies the key image features by the weight matrix to obtain the initial category score, which is a process of preliminarily quantifying the scenic subcategory and provides a benchmark for subsequent correction. Analyze the spatial distribution of the key image feature vectors in the environmental image and correct the score according to the result. When the key image feature vectors are concentratedly distributed, correct according to the first correction strategy, which can avoid misjudgment caused by excessive concentration of local features and highlight the influence of overall features on the scenic category; if they are dispersedly distributed, correct according to the second correction strategy, which can comprehensively consider the contribution of dispersed features, so that the category score can more truly reflect the overall scenic category tendency of the image, and further improve the accuracy and reliability of scenic category prediction.

[0083] Step S13: Predict the target scenic category hit by the environmental image based on the association relationship between the preset scenic subcategory and the scenic category and the category score corresponding to the scenic subcategory.

[0084] In the embodiment of the present application, step S13 includes the following steps C1 - C4:

[0085] Step C1: Obtain the association relationship between the preset scenic subcategory and the scenic category.

[0086] Specifically, based on the actual application requirements, the internal relationship between different sub - scenery categories and scenery categories is determined in advance, and a preset association relationship is established. This association relationship can be stored in the form of a specific data structure. For example, it can be in tabular form, creating a corresponding table structure in the database, where each row records a sub - scenery category and its corresponding scenery category; or it can be stored in a configuration file in a specific format (such as JSON format, etc.) to facilitate quick and accurate acquisition during the program operation.

[0087] As an example, as Figure 2 shown, European - style architecture, bridges, skyscrapers, and Ferris wheels, as sub - scenery categories, all belong to the scenery category of architecture; lakeside roads, mountain - ring roads, and tree - lined roads, as sub - scenery categories, all belong to the scenery category of roads; blue sky, sunset glow, and rainbow, as sub - scenery categories, all belong to the scenery category of weather.

[0088] Step C2: Determine the scenery category associated with each sub - scenery category according to the association relationship.

[0089] Specifically, based on the association relationship between the preset sub - scenery category and scenery category, when processing each sub - scenery category, its belonging scenery category can be determined by looking up the corresponding association mapping. For example, when dealing with the sub - scenery category of European - style architecture, it can be determined through querying the association relationship that it is associated with the scenery category of architecture. Similarly, for the sub - scenery category of lakeside road, it can be determined that it belongs to the scenery category of road.

[0090] Step C3: Calculate the probability prediction value of each scenery category corresponding to the environmental image based on the category score.

[0091] In the embodiment of the present application, Step C3 includes the following steps C31 - C33:

[0092] Step C31: Calculate the initial probability prediction value of each scenery category corresponding to the environmental image based on the category score.

[0093] Specifically, an activation function (such as Softmax) can be used to convert the category score into the initial probability prediction value of each scenery category. The formula of the Softmax function is:

[0094]

[0095] where P(y i ) represents the initial probability prediction value of scenery category i, and z i is the category score of sub - scenery category i. The Softmax function ensures that the sum of the probabilities of all scenery categories is 1, making the output interpretable as a probability distribution.

[0096] Step C32: Detect whether there is an association relationship between each scenic sub-category in the scenic category.

[0097] It should be noted that there may be internal connections or correlations between different scenic sub-categories, which will affect the determination of the final scenic category. For example, in the architecture category, European-style architecture usually has a matching relationship with specific road styles (such as classical-style streets) or the surrounding environment (such as parks, squares, etc.); in the weather category, the appearance of a rainbow is usually associated with specific sky conditions (such as a blue sky).

[0098] Specifically, when detecting the association relationship, an association relationship matrix or lookup table can be pre-constructed, which records the degree of association or rules between different scenic sub-categories. For example, if 0 represents no association and 1 represents association, then the corresponding element in the association relationship matrix for European-style architecture and classical streets is 1. The program will traverse this matrix or lookup table and check whether there is a set association relationship for the scenic sub-categories involved in the current environmental image. It is also possible to use a machine learning model trained based on a large amount of data to judge the association relationship between scenic sub-categories. The model can learn information such as the probability of different sub-categories co-occurring in the actual scene, so as to determine whether there is an association relationship between them, providing a basis for subsequent probability correction.

[0099] Step C33: If there is an association relationship, correct the initial probability prediction value of the environmental perception category according to the third correction strategy to obtain the probability prediction value.

[0100] Specifically, when scenic sub-categories with an association relationship are detected in the scenic category, it is necessary to correct the initial probability prediction value calculated previously. The core of the third correction strategy is to adjust the probability according to factors such as the tightness of the association relationship and the importance of the relevant scenic sub-categories in the image. For example, if both European-style architecture and classical streets, two associated scenic sub-categories, are detected in the image, and the category score of European-style architecture is relatively high, it indicates that the image is likely to be a scene with European-style architecture and street landscapes. Then, the initial probability prediction value for the architecture scenic category should be appropriately increased. An association correction coefficient k (k > 1, whose value is determined according to the strength of the association relationship and experience, such as k = 1.5 when the association is very tight) can be set, and the initial probability prediction value of the architecture scenic category is multiplied by k to obtain the corrected probability prediction value.

[0101] The method provided by the embodiment of the present application calculates the initial probability prediction value of each scenic category corresponding to the environmental image based on the category score, providing starting data for subsequent accurate prediction. Detecting whether there is an association relationship between each scenic sub-category in the scenic category can uncover the hidden connections between sub-categories. If there is an association relationship, the initial probability prediction value is corrected according to the third correction strategy to obtain the probability prediction value, which can fully consider the impact of sub-category association on the judgment of the scenic category, avoid misjudgment caused by treating sub-categories in isolation, optimize the probability prediction result, and improve the accuracy and rationality of the scenic category prediction.

[0102] Step C4, use the scenic category corresponding to the maximum value in the probability prediction value as the target scenic category.

[0103] Specifically, after calculating and correcting the probability prediction value of each scenic category in the environmental image through a series of previous steps, the final probability prediction value corresponding to each scenic category will be obtained. For example, for the three scenic categories of buildings, roads, and weather, there are corresponding probability prediction values: P B , P R , P W . At this time, by comparing the magnitudes of these probability prediction values, the maximum value can be found. Suppose P B > P R and P B > P W , then it can be determined that the scenic category of buildings is the target scenic category hit by this environmental image.

[0104] The method provided by the embodiment of the present application obtains the association relationship between the preset scenic sub-category and the scenic category and determines the scenic category associated with each scenic sub-category, which helps to construct a complete framework of the scenic classification system, enabling the model to clarify the attribution and association of different sub-categories in the large category. Calculating the probability prediction value of each scenic category corresponding to the environmental image based on the category score is to convert the sub-category information into a prediction basis at the category level. Finally, using the scenic category corresponding to the maximum value in the probability prediction value as the target scenic category can accurately screen out the scenic category that best conforms to the image features, improve the accuracy and certainty of scenic classification, and make the entire scenic recognition process more logical and effective.

[0105] Step S14, obtain the target environmental image associated with the target scenic category from the environmental image sequence, and generate a scenic video corresponding to the target scenic category using the target environmental image.

[0106] In the embodiment of the present application, generating a scenic video corresponding to the target scenic category using the target environmental image includes the following steps D1 - D3:

[0107] Step D1, detect the image quality of the target environmental image to obtain the quality score of the target environmental image.

[0108] Specifically, first, the detection of image quality takes into account multiple factors. For clarity, it can be evaluated by calculating the gradient information of the image. For example, the Sobel operator is used to calculate the gradient magnitude in the horizontal and vertical directions of the image. The larger the gradient magnitude, the clearer the object edges in the image and the higher the overall clarity. It is also possible to calculate the variance of the image. If the variance is large, it indicates that the pixel values of the image are more dispersed, and the image has high clarity and rich detail information. In terms of brightness and contrast, calculate the average brightness value of the image and the dynamic range of brightness (i.e., the difference between the brightest pixel value and the darkest pixel value in the image). Appropriate brightness and contrast can make the objects in the image easier to distinguish and recognize. For example, if the average brightness value is close to the middle gray value (such as 127 for an 8-bit grayscale image) and the dynamic range is large, it means that the brightness and contrast are more appropriate. Color accuracy is also one of the important factors. It can be compared with a standard color space (such as sRGB, etc.), and the deviation degree of the image color in different color channels (such as the RGB channel) is calculated. The smaller the deviation, the higher the color accuracy.

[0109] Combining these factors, a quality score is given to the target environment image by means of weighted summation or establishing a multi-factor evaluation model (such as an image quality evaluation model obtained by machine learning training). For example, it can be set that clarity accounts for 40%, brightness and contrast account for 40%, and color accuracy accounts for 20%. The final quality score is obtained by weighting according to the calculation results of each factor. The higher the score, the better the image quality. This score will be used as the basis for subsequent image screening.

[0110] Step D2, based on the quality score, screen the target environment image to obtain candidate environment images.

[0111] Specifically, after obtaining the quality score, it is necessary to determine the screening threshold or strategy. Among them, the screening threshold can be a fixed threshold, such as 70 points (out of 100 points). Target environment images with a score exceeding 70 points are used as candidate environment images. It is also possible to dynamically set the screening threshold. First, calculate the mean and standard deviation of all image scores. For example, the mean is 60 points and the standard deviation is 10 points. It is determined according to the mean plus several times the standard deviation (such as mean + 1×standard deviation = 70 points), and adjusted flexibly according to the actual situation to retain high-quality images. It is also possible to screen according to a ratio. For example, retain the top 50% of the target environment images with the highest scores as candidate environment images.

[0112] Step D3, match the corresponding video template according to the candidate environment image, and generate the corresponding landscape video based on the video template and the candidate environment image.

[0113] Specifically. A video template library with different styles and themes is pre-built. Each template has different characteristics in screen layout, transition, and music matching. For example, the architectural template highlights the majestic details of the building and the transition is smooth. The road template has a strong sense of dynamics and brisk music. The weather template is adjusted according to the weather color and music. Then match the corresponding template according to the target scenery category. For example, the building category matches the architectural template. You can also fine-tune according to the specific features of the candidate image. If the architectural style is classical, select an architectural template with classical elements and warm tones. Finally, synthesize the candidate image according to the selected template, embed the image according to its picture sequence, duration, transition, etc., add music subtitles to generate a complete landscape video, such as the architectural video displays European-style building images in layout order, fades in and out transitions, and is accompanied by classical music and architectural subtitle information for users to record their journey.

[0114] The method provided in the embodiment of the present application detects the image quality of the target environment image and obtains a quality score, which can screen out target environment images with higher image quality, eliminate the adverse effects of low-quality images such as blur and high noise on the generation of landscape videos, and screen out candidate environment images based on the quality score to ensure that the images used for video generation have certain quality standards. According to the candidate environment images, the corresponding video template is matched and the landscape video is generated, so that the style of the generated landscape video is consistent with the image content, the viewing quality and professionalism of the video are improved, and the user is provided with high-quality landscape video records and the user experience is enhanced.

[0115] In the embodiment of the present application, the category score corresponding to the scenery subcategory is calculated according to the input environment image sequence, and the category score corresponding to the scenery subcategory is converted into a probability prediction value of the scenery category, and the target scenery category corresponding to the maximum value of the probability prediction value is output. A pre-trained visual model can be used, and the model structure of the visual model is as follows: Figure 3 As shown, it includes: 7x7 convolution layer, 3x3 maximum pooling layer, residual block structure, global average pooling layer, and fully connected layer.

[0116] The 7x7 convolutional layer is used to extract the primary features of the image and use a convolution operation with a stride of 2 to reduce the spatial dimension of the image. It uses a convolution kernel with a width and height of 7 pixels to capture local features in the image while covering a larger receptive field, thereby extracting more extensive contextual information.

[0117] The 3x3 maximum pooling layer is used to reduce the spatial size of the feature map while retaining important information. It slides on the feature map using a 3x3 window, each window covers an adjacent 3x3 pixel area, moves the window according to the rule of a step size of 2 (that is, moves two pixels each time), and takes the maximum value of all elements in each window as the output, which helps the network obtain translation invariance and makes the network insensitive to small position changes of objects in the input image;

[0118] The residual block structure consists of four stages. Stage 1 includes 3 residual blocks, and each residual block is composed of multiple convolutional layers and possible shortcut connections. As the network depth increases, the convolutional neural network (CNN) can learn more complex feature hierarchies. The deep network can learn more abstract high-level features using this stage, and at the same time, use the shortcut connections to solve the problem of vanishing gradients. Stage 2 includes 4 residual blocks, similar to Stage 1, but this stage has more filters, which enables the network to capture more complex features and enrich the ability to extract deep features of the image. Stage 3 includes 6 residual blocks, further increasing the network depth and allowing the network to learn more abstract semantic feature representations. Stage 4 includes 3 residual blocks, marking the end of the network and preparing for outputting the target landscape category. Through these stages, the image features are mined and integrated step by step;

[0119] The global average pooling layer is used to reduce the spatial dimension of the feature map after being processed by multiple residual blocks to 1, preparing a suitable data format for the fully connected layer, so that the subsequent fully connected layer can perform classification-related operations based on the processed feature information;

[0120] The fully connected layer is used to linearly combine the vectors after global average pooling, calculate the scores of each subcategory. Based on the operations of the input vector and the weight matrix and the effect of the bias term, the scores corresponding to each category are obtained, and then the activation function is used to convert the category branches into probability prediction values, and then determine the final output target landscape category (that is, the category with the highest predicted probability value among the three landscape categories of building, road, and weather) to complete the output of the image classification task.

[0121] In the embodiments of this application, the specific implementation process includes:

[0122] The 7x7 convolutional layer is used to receive the environmental image sequence input to the visual model, perform convolutional processing on the environmental images in the environmental image sequence to obtain multiple image features, and transmit the image features to the 3x3 max pooling layer;

[0123] The 3x3 max pooling layer is used to receive the image features transmitted by the 7x7 convolutional layer, extract the maximum value in the window elements during the sliding process through a preset window to obtain the element maximum value, and transmit the element maximum value to the residual block structure;

[0124] The residual block structure is used to receive the element maximum value transmitted by the 3x3 max pooling layer, and sequentially process the element maximum value through the residual blocks of each stage to obtain the processed image features, and transmit the processed image features to the global average pooling layer;

[0125] A global average pooling layer, which is used to receive the processed image features transmitted by the residual block structure, reduce the spatial dimension of the processed image features to a specified dimension to obtain key image features, and transmit the key image features to a fully connected layer;

[0126] A fully connected layer, which is used to receive the key image features transmitted by the global average pooling layer, calculate the category scores of each sub-category of scenery through the key image features and a weight matrix, and convert the category scores of the sub-categories of scenery into predicted probability values of the scenery category by using a preset activation function. Determine the maximum value among the predicted probability values of each scenery category and output the target scenery category corresponding to the maximum value.

[0127] In the embodiments of the present application, as Figure 4 shown, the application process of the visual model includes: obtaining a plurality of environmental images collected during the driving of the target vehicle; predicting the target scenery category hit by the environmental images through the visual model; removing the environmental images with a similarity exceeding a preset similarity in the target scenery category through a deduplication model; removing the environmental images with a quality score not meeting the preset conditions through an aesthetic evaluation model; matching the corresponding video templates according to the reserved candidate environmental images; and generating a corresponding scenery video based on the video template and the candidate environmental images.

[0128] In this embodiment, a configuration method of a visual model is provided. Figure 5 It is a flowchart of the configuration method of the visual model according to the embodiments of the present invention. As Figure 5 shown, the process includes the following steps:

[0129] Step S21, obtaining the operation data of the target vehicle, where the operation data includes a plurality of environmental images.

[0130] In the embodiments of the present application, the environmental images in the operation data of the target vehicle are mainly obtained by various in-vehicle image acquisition devices, such as a driving recorder and cameras with different perspectives such as front view, rear view, and side view. The driving recorder can continuously record the road conditions in front of the vehicle. For example, when the vehicle is driving on an urban road, it can capture images of high-rise buildings (corresponding to the sub-category of scenery under the category of buildings - skyscrapers). The in-vehicle cameras with different perspectives record the surrounding environment of the vehicle in all directions. The side view camera can obtain images of vehicles in adjacent lanes, roadside greenery, and buildings. The rear view camera can obtain the road conditions and traffic conditions behind the vehicle. These images form the original set of environmental images for model training.

[0131] In addition, the training dataset can be extended by using environmental images in combination with data augmentation techniques. The data augmentation techniques include randomly rotating, scaling, translating, and flipping images, as well as color operations such as adjusting brightness, contrast, and saturation (the probability of horizontal flipping and random rotation within 0 to 90 degrees for each image is 50%). Gaussian noise can also be introduced, and affine transformation and perspective transformation can be performed to improve the accuracy of the model in recognizing targets under different viewpoints, lighting conditions, and occlusions.

[0132] Step S22: Perform quantization-aware training on the original visual model based on the operation data to obtain the trained original visual model.

[0133] In the embodiment of the present application, first, preprocess the environmental images in the obtained operation data, such as unifying the resolution size and adjusting the color mode to make them conform to the input data specification of the original visual model. Then, use the quantization-aware training (QAT) technique to train the original visual model, which can determine the quantization parameters required to convert the floating-point (float32) parameters in the model into low-precision integer (int8) representations (calculated by the formula output=(input - offset)*scale, where offset and scale are obtained through quantization-aware training). During training, the model learns based on the preprocessed environmental images and the corresponding annotation information (such as target category annotations for buildings, roads, weather, etc.). As the training iterates, its internal weights and other parameters are continuously optimized, enabling the model to not only master the different target features in the environmental images but also use quantization-aware training to reduce the computational complexity and memory occupancy. While maintaining the prediction accuracy for various targets, it is more suitable for the in-vehicle environment. Finally, the obtained trained original visual model is significantly improved in terms of accuracy and resource utilization efficiency and can better meet the subsequent actual application requirements.

[0134] Step S23: Convert the trained original visual model into a configuration file corresponding to the visual model suitable for a specific file format.

[0135] In the embodiment of the present application, after completing the quantization-aware training to obtain the trained original visual model, to successfully deploy and efficiently operate it in the in-vehicle environment of the target vehicle, it needs to be converted into a specific file format and the corresponding configuration file is generated. Use the model conversion tool of Qualcomm Neural Network (QNN) to convert. First, convert the trained original visual model into.cpp and.bin formats, which are the basic formats for running in the in-vehicle environment. Then, further compile it in combination with the.json configuration file containing key information such as the model structure, parameters of each layer, and data flow direction to accurately convert the model into a binary file that meets the requirements of the in-vehicle environment, that is, the configuration file corresponding to the visual model suitable for a specific file format.

[0136] Step S24: Install the configuration file into the target vehicle, so that the target vehicle processes the environmental image sequence during driving by using the vision model to obtain a scenic video.

[0137] In the embodiment of the present application, first, the configuration file is transmitted and installed into the corresponding system of the target vehicle through a suitable method. Specifically, it can be copied wired by using the on-board diagnostic interface (OBD), or remotely installed wirelessly by means of the built-in wireless communication module of the vehicle (such as a 4G / 5G module). After successful installation, when the vehicle is driving, the vision model carried by it can process the environmental image sequence collected by devices such as cameras in real time. The vision model analyzes each image according to the learned features and classification capabilities, determines the target category to which the element belongs (such as buildings, roads, weather, etc.), and then screens and integrates the image information with classification results according to the recognition situation, preset rules, and algorithms. For example, after identifying the images of the scenic category, these images are combined logically in a certain order, and combined with video editing strategies (such as adding transition effects, matching music, etc.) to generate an ornamental scenic video, which is used to record the wonderful scenery along the way, provides a way to record the journey for the driver and passengers, and reflects the value of the vision model in the in-vehicle application scenario.

[0138] The method provided by the embodiment of the present application can collect diverse image information by obtaining the operation data of the target vehicle containing multiple environmental images, builds a comprehensive material library for the model, enhances the generalization ability of the model, and helps it accurately identify targets in the in-vehicle environment. Based on the operation data, performing quantization perception training on the original vision model can convert the quantization parameters required for parameter conversion, enable the model to learn the subtle features of the target to improve the recognition accuracy, and can also reduce the computational complexity and memory occupancy, optimize the performance and adapt to the in-vehicle environment. Converting the trained original vision model into a configuration file corresponding to a specific file format solves the compatibility problem between the model and the in-vehicle software and hardware environment, converts the format and integrates key information, ensures smooth deployment and stable operation, and lays a solid foundation for the application. Installing the configuration file into the target vehicle, when the vehicle is driving, using the vision model to process the environmental image sequence to generate a scenic video, which is convenient for the driver and passengers to record the journey scenery, enhances the intelligence and fun of the vehicle, realizes the retention of scenic image materials, and improves the user experience in the in-vehicle scenario.

[0139] In this embodiment, a device for generating a scenic video is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0140] This embodiment provides a device for generating a scenic video, such asFigure 6 As shown, it includes:

[0141] An acquisition module 61, configured to acquire an environmental image sequence collected by a target vehicle during driving;

[0142] An identification module 62, configured to identify key image features of each environmental image in the environmental image sequence, and obtain category scores of the landscape subcategories corresponding to the key image features;

[0143] A prediction module 63, configured to predict a target landscape category hit by the environmental image based on the association relationship between preset landscape subcategories and landscape categories and the category scores corresponding to the landscape subcategories;

[0144] A generation module 64, configured to obtain a target environmental image associated with the target landscape category from the environmental image sequence, and generate a landscape video corresponding to the target landscape category by using the target environmental image.

[0145] In an alternative embodiment of the present application, the identification module 62 includes an extraction sub-module and a calculation sub-module;

[0146] The extraction sub-module is configured to identify original image features in the environmental image; use a sliding window to extract the element maximum value in the original image features until the spatial dimension of the original image features in the environmental image reaches a specified dimension, and take the element maximum value as the key image feature.

[0147] The calculation sub-module is configured to obtain a weight matrix corresponding to the landscape subcategory; calculate the category score of the landscape subcategory corresponding to the key image feature according to the key image feature and the weight matrix.

[0148] The calculation sub-module is configured to multiply the key image feature by the weight matrix to obtain an initial category score corresponding to the landscape subcategory; analyze the spatial distribution of the key image feature vectors in the environmental image to obtain an analysis result; if the analysis result is that the key image feature vectors are concentratedly distributed, then correct the initial category score of the landscape subcategory according to the first correction strategy to obtain the category score; or, if the analysis result is that the key image feature vectors are dispersedly distributed, then correct the initial category score of the landscape subcategory according to the second correction strategy to obtain the category score, where the first correction strategy is used to perform a correction operation on the initial category score according to the score improvement coefficient corresponding to the concentrated distribution, and the second correction strategy is used to perform a correction operation on the initial category score according to the score reduction coefficient corresponding to the dispersed distribution.

[0149] In an alternative embodiment of the present application, the prediction module 63 is configured to obtain the association relationship between the preset scenic sub - categories and the scenic categories; determine the scenic category associated with each scenic sub - category according to the association relationship; calculate the probability prediction value of each scenic category corresponding to the environmental image based on the category score; and use the scenic category corresponding to the maximum value in the probability prediction values as the target scenic category.

[0150] In an alternative embodiment of the present application, the prediction module 63 is configured to calculate the initial probability prediction value of each scenic category corresponding to the environmental image based on the category score; detect whether there is an association relationship between each scenic sub - category in the scenic category; if there is an association relationship, correct the initial probability prediction value of the environmental perception category according to the third correction strategy to obtain the probability prediction value, where the third correction strategy is used to perform a correction operation on the initial probability prediction value according to the tightness of the association relationship and the importance of the scenic sub - category in the environmental image.

[0151] In an alternative embodiment of the present application, the generation module 64 is configured to detect the image quality of the target environmental image to obtain the quality score of the target environmental image; screen the target environmental image based on the quality score to obtain candidate environmental images; match the candidate environmental images with the corresponding video templates, and generate the corresponding scenic video based on the video templates and the candidate environmental images.

[0152] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 7 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high - speed interface and a low - speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi - processor system).

[0153] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0154] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.

[0155] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of a computer device presented by a kind of mini-program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided relative to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0156] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.

[0157] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.

[0158] Embodiments of the present invention also provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading over a network from an original storage in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the methods described herein can be stored as such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0159] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for generating a landscape video, characterized in that, The method includes: Obtaining an environmental image sequence collected during the driving of a target vehicle; Identifying key image features of each environmental image in the environmental image sequence, and obtaining the category scores of the scenic sub-categories corresponding to the key image features; Predicting the target scenic category hit by the environmental image based on the preset association relationship between the scenic sub-categories and the scenic categories and the category scores corresponding to the scenic sub-categories; Obtaining the target environmental image associated with the target scenic category from the environmental image sequence, and generating a scenic video corresponding to the target scenic category by using the target environmental image.

2. The method according to claim 1, wherein The identifying the key image features in the environmental image includes: Identifying the original image features in the environmental image; Using a sliding window to extract the element maximum value in the original image features until the spatial dimension of the original image features in the environmental image reaches the specified dimension, and taking the element maximum value as the key image feature.

3. The method according to claim 1, wherein The obtaining the category scores of the scenic sub-categories corresponding to the key image features includes: Obtaining the weight matrix corresponding to the scenic sub-category; Calculating the category scores of the scenic sub-categories corresponding to the key image features according to the key image features and the weight matrix.

4. The method according to claim 3, wherein The calculating the category scores of the scenic sub-categories corresponding to the key image features according to the key image features and the weight matrix includes: Multiplying the key image features by the weight matrix to obtain the initial category scores corresponding to the scenic sub-categories; Analyzing the spatial distribution of the key image feature vectors in the environmental image to obtain an analysis result; If the analysis result is that the key image feature vectors are concentrated, then correcting the initial category scores of the scenic sub-categories according to the first correction strategy to obtain the category scores; or, if the analysis result is that the key image feature vectors are dispersed, then correcting the initial category scores of the scenic sub-categories according to the second correction strategy to obtain the category scores, where the first correction strategy is used to perform a correction operation on the initial category scores according to the score improvement coefficient corresponding to the concentrated distribution, and the second correction strategy is used to perform a correction operation on the initial category scores according to the score reduction coefficient corresponding to the dispersed distribution.

5. The method according to claim 1, wherein The predicting the target scenic category hit by the environmental image based on the preset association relationship between the scenic sub-categories and the scenic categories and the category scores corresponding to the scenic sub-categories includes: Obtaining the association relationship between the preset scenic sub-categories and the scenic categories; Determining the scenic categories associated with each of the scenic sub-categories according to the association relationship; Calculating the probability prediction values of each scenic category corresponding to the environmental image based on the category scores; Taking the scenic category corresponding to the maximum value in the probability prediction values as the target scenic category.

6. The method according to claim 5, wherein The calculating the probability prediction values of each scenic category corresponding to the environmental image based on the category scores includes: Calculating the initial probability prediction values of each scenic category corresponding to the environmental image based on the category scores; Detecting whether there is an association relationship between each of the scenic sub-categories in the scenic category; If there is an association relationship, the initial probability prediction value of the environmental perception category is corrected according to the third correction strategy to obtain a probability prediction value, where the third correction strategy is used to perform a correction operation on the initial probability prediction value according to the tightness of the association relationship and the importance of the landscape sub-category in the environmental image.

7. The method according to claim 1, wherein The generating the landscape video corresponding to the target landscape category by using the target environmental image includes: Detecting the image quality of the target environmental image to obtain a quality score of the target environmental image; Screening the target environmental image based on the quality score to obtain a candidate environmental image; Matching the candidate environmental image with a corresponding video template, and generating a corresponding landscape video based on the video template and the candidate environmental image.

8. An apparatus for generating a scenic video, characterized in that, The device includes: An acquisition module, configured to acquire an environmental image sequence collected during the driving of a target vehicle; An identification module, configured to identify key image features of each environmental image in the environmental image sequence, and obtain category scores of the landscape sub-categories corresponding to the key image features; A prediction module, configured to predict a target landscape category hit by the environmental image based on a preset association relationship between the landscape sub-category and the landscape category and the category score corresponding to the landscape sub-category; A generation module, configured to obtain a target environmental image associated with the target landscape category from the environmental image sequence, and generate a landscape video corresponding to the target landscape category by using the target environmental image.

9. A computer device, characterized in that, including: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.