Vehicle multimedia data synthesis method and device, computer equipment and medium
By acquiring and integrating environmental images during driving in the vehicle, building an aggregated image group and synthesizing images according to user needs, the problem of fragmented and difficult image acquisition and integration is solved, and the multimedia data generation with logical coherence and orderly content is achieved, meeting users' depth and integrity needs for travel memories.
Patent Information
- Application Number
- CN202510290492.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-10
AI Technical Summary
Due to fragmented pictures and driving fatigue, the beautiful scenery is missed, and the pictures lack systematic integration, making it difficult for users to connect the scenery and events of the journey, resulting in the fragmentation of memories.
By obtaining the original environment images during the vehicle's driving process, aggregating image groups, aggregating similar environment images, and synthesizing images according to user style needs to generate logically coherent and ordered multimedia data.
The preliminary classification and integration of a large number of fragmented pictures has been achieved, the efficiency and accuracy of data processing have been improved, and the generated multimedia data is logically coherent and the content is orderly, meeting the users' needs for depth and integrity of travel memories.
Smart Images

Figure CN120128809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle multimedia, and particularly to a method, device, computer device and medium for synthesizing vehicle multimedia data. Background Art
[0002] With the widespread use of automobiles, people's travel has become more convenient and frequent. In travel scenarios, people tend to record moments through pictures and share them in online communities to meet their spiritual needs. As an important carrier for information dissemination in online communities, the richness, uniqueness, and impression depth of the content of pictures determine the information transmission effect. Currently, pictures are usually captured by mobile phones or cameras. However, due to factors such as shooting scenes, time, and personal attention, the captured pictures are rather fragmented. As time goes by, in the face of a large number of fragmented pictures, it is difficult for users to connect the scenery and events during a certain journey. In addition, as the number of recorded events increases and the driving time lengthens, the human body is prone to fatigue. This not only affects the driving experience but also causes the driver to miss many beautiful sceneries or specific events along the way. At the same time, due to the lack of systematic integration of a large number of fragmented pictures, users' memories of the journey tend to be fragmented.
[0003] Although there are currently some general picture processing technologies, such as grouping processing (fragmented pictures are grouped and segmented according to certain rules, and users need to click on each group to view the internal pictures) and stitching / synthesizing pictures (with the help of stitching / synthesizing algorithms, users manually upload pictures and place them according to preset fixed positions, and then stitch or synthesize a new picture). However, these methods are not ideal in dealing with the problem of picture fragmentation caused by the increase in the number of trips and the passage of time. As time goes by, users' impressions of a certain journey will still gradually fade, making it difficult to form a complete, profound, and lasting memory, and unable to meet users' needs for the depth and integrity of travel memories. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, device, computer device and medium for synthesizing vehicle multimedia data to solve the problem that it is difficult for users to connect the scenery and events during the journey due to fragmented picture acquisition, missing beautiful sceneries caused by driving fatigue, and lack of systematic integration of pictures.
[0005] In a first aspect, an embodiment of the present invention provides a method for synthesizing vehicle multimedia data, the method comprising:
[0006] Obtaining a plurality of original environmental images;
[0007] Constructing a plurality of aggregated image groups by using the original environmental images, wherein the plurality of original environmental images in each aggregated image group are similar environmental images;
[0008] Aggregate the original environmental images in each of the aggregated image groups to obtain an aggregated environmental image corresponding to each aggregated image group;
[0009] Obtain the style requirements input by the user, and process the aggregated environmental image using the image synthesis strategy corresponding to the style requirements to generate the multimedia data corresponding to the style requirements.
[0010] Further, the constructing of multiple aggregated image groups using the original environmental images includes:
[0011] Calculate the eigenvalue vector of each of the original environmental images;
[0012] Based on the eigenvalue vector, determine the environmental images with similar features among the multiple original environmental images, and use the environmental images with similar features to construct multiple aggregated image groups.
[0013] Further, the calculating of the eigenvalue vector of each of the original environmental images includes:
[0014] = Extract the pixel information in the original environmental image;
[0015] Determine whether the original environmental image is a valid image according to the pixel information;
[0016] When the original environmental image is a valid image, extract the image features under different metrics in the original environmental image;
[0017] Use the image features under different metrics in the original environmental image to construct the eigenvalue vector.
[0018] Further, the determining of the environmental images with similar features among the multiple original environmental images based on the eigenvalue vector, and using the environmental images with similar features to construct multiple aggregated image groups includes:
[0019] Select a first environmental image from the original environmental images, and construct a first aggregated image group based on the first environmental image;
[0020] Use the eigenvalue vector to select a second environmental image similar to the first environmental image from the multiple original environmental images, and assign the second environmental image to the first aggregated image group;
[0021] Select a third environmental image from the remaining environmental images, and construct a second aggregated image group based on the third environmental image, where the remaining environmental images are the environmental images in the original environmental images except the first environmental image and the second environmental image;
[0022] Select a fourth environmental image similar to the third environmental image from the multiple remaining environmental images by using the eigenvalue vector, and allocate the fourth environmental image to the second aggregated image group;
[0023] Repeat the image allocation operation until all the original environmental images are allocated, and obtain multiple aggregated image groups.
[0024] Further, the step of selecting a second environmental image similar to the first environmental image from the multiple original environmental images by using the eigenvalue vector includes:
[0025] Calculate the similarity between the eigenvalue vector of the first environmental image and the eigenvalue vectors of the other environmental images except the first environmental image among the original environmental images to obtain a similarity matrix;
[0026] Determine at least one second environmental image similar to the first environmental image among the other environmental images according to the similarity matrix.
[0027] Further, the step of processing the aggregated environmental image by using the image synthesis strategy corresponding to the style requirement to generate the multimedia data corresponding to the style requirement includes:
[0028] Determine the target features and the image synthesis strategy according to the style requirement;
[0029] Analyze the matching degree between the eigenvalue vector of each aggregated environmental image and the target features to obtain a matching result;
[0030] Determine the target environmental image among the aggregated environmental images according to the matching result, and process the target environmental image by using the image synthesis strategy to generate the multimedia data corresponding to the style requirement.
[0031] Further, the step of processing the target environmental image by using the image synthesis strategy to generate the multimedia data corresponding to the style requirement includes:
[0032] Determine the corresponding data type according to the eigenvalue vector of the target environmental image;
[0033] Obtain the layout mode associated with the data type, and mark each target environmental image by using the time arrangement rule to obtain the target environmental image carrying the serial number identifier;
[0034] Process the target environmental image carrying the serial number identifier according to the layout mode to generate the multimedia data corresponding to the style requirement.
[0035] Further, processing the target environmental image carrying the serial number identifier according to the layout pattern to generate the multimedia data corresponding to the style requirement includes:
[0036] If the data type is a regular type, adjust the target environmental image carrying the serial number identifier according to the layout parameters in the first layout pattern, and sequentially fill the adjusted target environmental image into the designated positions of the first layout pattern according to the serial number identifier to obtain multimedia data;
[0037] If the data type is an irregular type, analyze the positional relationship between the target environmental images carrying the serial number identifier according to the second layout pattern, and fuse the target environmental images based on the positional relationship and the serial number identifier to obtain multimedia data.
[0038] In a second aspect, an embodiment of the present invention provides a vehicle multimedia data synthesis device, which includes:
[0039] An acquisition module, configured to acquire a plurality of original environmental images;
[0040] A construction module, configured to construct a plurality of aggregated image groups by using the original environmental images, wherein the plurality of original environmental images in each aggregated image group are similar environmental images;
[0041] An aggregation module, configured to aggregate the original environmental images in each aggregated image group to obtain an aggregated environmental image corresponding to each aggregated image group;
[0042] A processing module, configured to obtain a style requirement input by a user, and process the aggregated environmental image by using an image synthesis strategy corresponding to the style requirement to generate multimedia data corresponding to the style requirement.
[0043] In a third aspect, an embodiment of the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any corresponding embodiment thereof.
[0044] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method according to the first aspect or any corresponding embodiment thereof.
[0045] The method provided by the embodiments of the present application has the following beneficial effects:
[0046] The method provided by the embodiments of this application obtains the original environmental images during the vehicle's driving process, providing a rich data basis for subsequent synthesis of comprehensive and macroscopic multimedia data, ensuring that the information of all stages and scenes during the journey can be covered, and avoiding the inability to fully present the journey due to data loss. Aggregating similar environmental images helps to initially classify and integrate a large number of fragmented original environmental images, making subsequent processing more targeted, improving the efficiency and accuracy of data processing, facilitating the subsequent generation of logically coherent and content-ordered multimedia data, and helping users better connect the scenery and events of the journey. Aggregating the images within the aggregated image group further integrates the image information of similar scenes, reduces data redundancy, enhances the conciseness of image information, provides more optimized materials for generating multimedia data that meets user needs, and helps to form a complete and ordered travel presentation. According to the expectations of different users for the presentation method of travel memories, corresponding image synthesis strategies are adopted to process the aggregated environmental images, generating multimedia data that meets the specific style requirements of users, meeting the personalized needs of users, enhancing the depth and integrity experience of users' travel memories, and enabling users to review the journey in the way they expect. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 is a flowchart of the method for synthesizing vehicle multimedia data according to an embodiment of the present invention;
[0049] Figure 2 is a flowchart of the grouping and aggregation of environmental image processing according to an embodiment of the present invention;
[0050] Figure 3 is a flowchart of the environmental image screening and multimedia data generation according to an embodiment of the present invention;
[0051] Figure 4 is a schematic diagram of the architecture of the vehicle environmental image acquisition and processing system according to an embodiment of the present invention;
[0052] Figure 5 is a working flowchart of the vehicle environmental image acquisition and processing system according to an embodiment of the present invention;
[0053] Figure 6 is a structural block diagram of the device for synthesizing vehicle multimedia data according to an embodiment of the present invention;
[0054] Figure 7 It is a schematic diagram of the hardware structure of the computer device according to an embodiment of the present invention. Detailed implementation manners
[0055] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] According to an embodiment of the present invention, there are provided a method, an apparatus, a computer device and a medium for synthesizing vehicle multimedia data. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0057] In this embodiment, a method for synthesizing vehicle multimedia data is provided. Figure 1 It is a flowchart of the method for synthesizing vehicle multimedia data according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:
[0058] Step S11, obtaining a plurality of original environmental images.
[0059] In the embodiment of the present application, a plurality of original environmental images can be realized through a trigger acquisition module in the target vehicle. Among them, the trigger acquisition module is a front-end module, and its specific implementation process includes: the user sets parameters such as light, color temperature, and brightness according to needs and adds them to the trigger; the trigger monitors the changes of external environmental parameters in real time, and when the parameters reach the set range, the trigger device is started; the start signal is transmitted to the body controller through the bus to drive the relevant cameras to collect photos; the acquisition module transmits the obtained original images to the picture acquisition algorithm module, and after preliminary processing and data extraction, directly outputs a plurality of original environmental images that meet the requirements, and finally stores the picture information in the form of eigenvalue vectors.
[0060] Step S12, constructing a plurality of aggregated image groups by using the original environmental images, where the plurality of original environmental images in each aggregated image group are similar environmental images.
[0061] In the embodiment of the present application, step S12 includes the following steps A1-A2:
[0062] Step A1, calculating the eigenvalue vector of each original environmental image.
[0063] In the embodiment of the present application, step A1 includes the following steps A11 - A14:
[0064] Step A11, extract the pixel information in the original environmental image.
[0065] Specifically, a pixel is the smallest unit that makes up an image, and the visual information of the image is stored in the form of pixels. Analyze the original environmental image collected during the driving of the target vehicle, and through specific algorithms and technical means, read the relevant information such as the color and brightness of each pixel point in the image. These pixel information are the basis for further analyzing and processing the image, and subsequent various judgments and operations on the image rely on the accurately extracted pixel information. For example, for a color photo, it is necessary to extract the red, green, and blue (RGB) color component values of each pixel point, and these values determine the color presented by the pixel in the image.
[0066] Step A12, determine whether the original environmental image is a valid image according to the pixel information.
[0067] Specifically, after obtaining the pixel information of the original environmental image, it can be determined whether the image is a valid image based on preset determination conditions. The determination conditions can revolve around aspects such as the pixel level of the image, whether there are incoherent phenomena such as local blurring or jumping in the picture, etc. If the pixel arrangement of the image is chaotic and does not conform to the pixel values of a normal image, or there are obvious local blurred areas in the picture, resulting in unclear visual information, or there is picture jumping, making the image content lack coherence, these all indicate that there is a problem with the image and it does not meet the standard of a valid image. For example, when the pixel color in a part of the image suddenly changes abnormally, forming a strong contrast with the surrounding pixels, and this change is not the content that the image itself wants to express, the image can be determined to be an invalid image. And for an image with neat pixel arrangement and clear and coherent picture, it is determined to be a valid image.
[0068] Step A13, when the original environmental image is a valid image, extract the image features under different indexes in the original environmental image.
[0069] Specifically, after determining that the original environmental image is a valid image, deeply analyze the image from multiple different index dimensions to extract its unique image features. These indexes can include but are not limited to the texture features, shape features, color distribution features, etc. of the image. For example, the texture features can reflect the texture structure of the image surface, such as whether it is smooth or rough, and the density of the texture; the shape features focus on the outline and geometric shape of the objects in the image; the color distribution features focus on analyzing the distribution and proportional relationship of different colors in the image. For example, for a landscape image, the blue tone distribution feature of the sky part, the shape feature of the mountains, and the texture feature of the trees can be extracted.
[0070] Step A14, construct an eigenvalue vector using the image features under different metrics in the original environmental image.
[0071] Specifically, after extracting the image features under different metrics of the original environmental image, these features are integrated to construct an eigenvalue vector. An eigenvalue vector is a mathematical representation form that quantifies and combines various features of an image in the form of a vector. The vector elements of each dimension correspond to a specific image feature metric. In this way, the complex features of the image are transformed into a numerical vector that is convenient for computer processing and analysis. For example, if the features of the three main metrics of the texture, shape, and color distribution of the image are extracted previously, the texture feature can be quantified as the value of the first dimension of the vector, the shape feature as the value of the second dimension, and the color distribution feature as the value of the third dimension, and so on, to form a multi-dimensional eigenvalue vector.
[0072] Step A2, based on the eigenvalue vector, determine the environmental images with similar features in multiple original environmental images, and use the environmental images with similar features to construct multiple aggregated image groups.
[0073] In the embodiment of the present application, step A2 includes the following steps A21 - A25:
[0074] Step A21, select a first environmental image from the original environmental images, and construct a first aggregated image group based on the first environmental image.
[0075] Specifically, after obtaining multiple original environmental images during the driving of the target vehicle, in order to classify the images with similar features, it is necessary to first select one as the first environmental image. This selection method can be random extraction, or it can be achieved according to preset initial conditions (such as selecting the first image according to the acquisition time sequence, or screening the most representative image based on basic features such as brightness and color richness).
[0076] Create a first aggregated image group based on the selected first environmental image. This group initially only contains the first environmental image. Subsequently, other original environmental images with similar features will be gradually added, and more aggregated image groups will be established based on the remaining images. Through iterative classification, the similar feature division of all original environmental images is finally achieved. This process lays the foundation for subsequent multimedia data generation (such as panoramic image synthesis or continuous video production), ensuring the visual coherence of the images in the same group and improving the fusion naturalness and overall performance.
[0077] Step A22, use the eigenvalue vector to select a second environmental image similar to the first environmental image from multiple original environmental images, and assign the second environmental image to the first aggregated image group.
[0078] Specifically, when determining the environmental images similar to the first environmental image, i.e., the second environmental images, these images are added to the first aggregated image group constructed based on the first environmental image. In this way, the first aggregated image group begins to gradually converge environmental images with similar features. As this process is continuously repeated, the first aggregated image group will be continuously enriched and finally form a set containing multiple similar environmental images. These similar images can be better integrated together in subsequent processing, such as generating a panoramic image or a continuous video, presenting a more coherent and holistic visual effect to the user, helping to solve the problem in the prior art that pictures are fragmented and difficult to integrate, and improving the user's satisfaction with the recording and display of travel images.
[0079] In the embodiment of the present application, step A22 includes the following steps A221 - A222:
[0080] Step A221, calculate the similarity between the eigenvalue vector of the first environmental image and the eigenvalue vectors of other environmental images in the original environmental images except the first environmental image, and obtain a similarity matrix.
[0081] Specifically, in the process of constructing the aggregated image group, the first aggregated image group has been created based on the first environmental image. To determine the similarity degree between other environmental images and the first environmental image, it is necessary to compare their eigenvalue vectors. The eigenvalue vector is constructed by extracting pixel information from the original environmental image, judging its validity, and extracting image features of different metrics, which quantifies various features of the image.
[0082] Use a specific similarity calculation algorithm, such as the cosine similarity algorithm, which measures the similarity between two vectors by calculating the cosine value of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are, that is, the more similar the corresponding image features are; such as the Euclidean distance algorithm, which calculates the straight-line distance between two vectors in a multi-dimensional space. The shorter the distance, the more similar the image features are.
[0083] For each environmental image in the original environmental images except the first environmental image, its eigenvalue vector is calculated for similarity with the eigenvalue vector of the first environmental image. Arrange these calculated similarity values in a certain order to form a matrix, that is, the similarity matrix. For example, if there are n images (except the first environmental image), the size of the similarity matrix is (n×1), and each element in the matrix represents the similarity between a certain environmental image and the first environmental image. This similarity matrix provides a quantitative basis for subsequent determination of which environmental images are similar to the first environmental image.
[0084] Step A222, determine at least one second environmental image similar to the first environmental image from other environmental images according to the similarity matrix.
[0085] Specifically, after obtaining the similarity matrix, it is necessary to screen out the environmental images similar to the first environmental image according to this matrix. A similarity threshold can be set, and this threshold can be adjusted according to the actual situation. Traverse each element in the similarity matrix and compare each element (i.e., the similarity between each environmental image and the first environmental image) with the set threshold. When the similarity between a certain environmental image and the first environmental image is greater than or equal to the threshold, this environmental image is determined as the second environmental image, that is, it is considered that this image has similar features to the first environmental image. In this way, those images that are relatively similar to the first environmental image in terms of features are found from numerous environmental images, preparing for the subsequent assignment of similar images to the first aggregated image group.
[0086] Step A23: Select a third environmental image from the remaining environmental images and construct a second aggregated image group based on the third environmental image, where the remaining environmental images are the environmental images in the original environmental images except the first environmental image and the second environmental image.
[0087] Specifically, after the first and second environmental images are assigned to the first aggregated image group, it is necessary to select a third environmental image from the remaining unassigned original environmental images (i.e., the images that have not been classified into any aggregated group). This selection method can follow the screening logic (such as random extraction or selection of representative images based on brightness and color features), and use the same mechanism to construct the second aggregated image group: taking the third environmental image as the first image in the group as the reference, and defining a new starting point for classifying similar features.
[0088] Step A24: Select a fourth environmental image similar to the third environmental image from multiple remaining environmental images using the eigenvalue vector and assign the fourth environmental image to the second aggregated image group.
[0089] Specifically, for the construction of the second aggregated image group, repeat the feature similarity matching process: calculate the similarity matrix between the eigenvalue vector of the third environmental image and the eigenvalue vectors of the remaining environmental images, and screen out the fourth environmental images that meet the requirements (such as similarity ≥ threshold) by setting a threshold. Add the screened fourth environmental images to the second aggregated image group to ensure the consistency of the image features within the group. This process acts on the new grouping reference (the third environmental image) and the remaining image subset.
[0090] Step A25: Repeat the image assignment operation until all the original environmental images are assigned, and multiple aggregated image groups are obtained.
[0091] Specifically, all classifications are completed through a cyclic iteration mechanism: each round of operation starts with a new reference image (such as the third and fifth environmental images) selected from the remaining environmental images, constructs a new aggregation group and assigns similar images until there are no remaining images to assign. Finally, multiple aggregation image groups are generated, and the image features in each group are highly similar. This process ensures that all original environmental images are reasonably classified.
[0092] Step S13, aggregating the original environment images in each aggregated image group to obtain an aggregated environment image corresponding to each aggregated image group.
[0093] In the embodiments of the present application, the aggregation operation is not a simple splicing or superposition, but a comprehensive processing of the original environment images in the aggregated image group using specific algorithms and techniques based on the image feature value model. For example, using the image synthesis algorithm, considering the similar features of each image, the color, contrast, brightness and other parameters of the image can be uniformly adjusted to make the images in the group more visually coordinated. At the same time, intelligent fusion operations can be performed based on the content and composition of the image to remove some repeated or redundant information and retain key features, thereby improving the overall quality and clarity of the image.
[0094] Taking a group of aggregated images containing similar scenery as an example, during the aggregation process, the main parts of the scenery in each picture, such as mountains, rivers, etc., are automatically identified, and then these main parts are optimized and combined so that the generated aggregated environmental image can more perfectly show the overall picture of this scenic area without the incoordination or information clutter caused by simple splicing.
[0095] Through such aggregation operations, each aggregated image group can generate a corresponding aggregated environment image. These aggregated environment images not only retain the similar features of the images in the original image group, but also improve the quality and coherence of the images through integration, providing high-quality materials for subsequent further processing according to the style requirements input by the user.
[0096] As an example, Figure 2 As shown, the process of environmental image processing. First, the fragmented original environmental images include original environmental image 1, original environmental image 2, etc. After being processed, these original environmental images are grouped to form multiple aggregated image groups, namely aggregated image group 1, aggregated image group 2, and finally aggregated image group m. Each aggregated image group contains several environmental images. For example, aggregated image group 1 contains the first environmental image, the second environmental image *x, etc. After further processing, these images generate corresponding aggregated environmental images, such as aggregated environmental image 1, aggregated environmental image 2, and finally aggregated environmental image m. The whole process shows the process from fragmented original environmental images to the formation of multiple aggregated environmental images, reflecting the grouping and aggregation processing of images.
[0097] Step S14: Obtain the style requirements input by the user, and use the image synthesis strategy corresponding to the style requirements to process the aggregated environmental image to generate multimedia data corresponding to the style requirements.
[0098] It should be noted that the style requirements input by the user are the key guidelines for the entire image processing process, which determine the presentation effect of the finally generated multimedia data. For example, if the user sets it to "retro style", then according to the existing knowledge in the image feature value model, the target features related to the retro style are determined. These target features may include specific color combinations (such as warm tones, yellowish-brown tones), the adjustment range of contrast (possibly lower to create a soft visual effect), and picture graininess, etc.
[0099] As an example, as Figure 3 shown, the process from screening the target environmental image to generating multimedia data. First, in the link of screening the target environmental image, a series of target environmental images will be obtained, such as target environmental image 1, target environmental image 2, etc. These target environmental images will be processed according to the style requirements input by the user. Finally, in the link of generating multimedia data, the processed target environmental images (such as target environmental image 1, target environmental image 2, target environmental image 4, target environmental image 5, etc.) are combined to generate multimedia data (such as a four-grid). The whole process reflects the process of screening and processing the target environmental image according to the user's style requirements and finally generating multimedia data.
[0100] In the embodiment of the present application, using the image synthesis strategy corresponding to the style requirements to process the aggregated environmental image to generate multimedia data corresponding to the style requirements includes the following steps B1 - B3:
[0101] Step B1: Determine the target features and the image synthesis strategy according to the style requirements.
[0102] Specifically, first analyze the style requirements input by the user. For example, if the user sets it to "retro style", then according to the existing knowledge in the image feature value model, the target features related to the retro style are determined. These target features may include specific color combinations (such as warm tones, yellowish-brown tones), the adjustment range of contrast (possibly lower to create a soft visual effect), and picture graininess, etc.
[0103] After determining the target features, select an image synthesis strategy that matches them. The image synthesis strategy is a series of operation steps designed based on computer algorithms and artificial intelligence technologies for different style requirements. For the retro style, the synthesis strategy can include adjusting the color balance of the image to make it approach the retro color tone, reducing the clarity of the image to simulate the texture of old photos, or adding some elements such as retro-style borders. These strategies are derived from training on a large amount of image data and analysis and summary of various style characteristics, aiming to accurately process the aggregated environmental images into multimedia data that meets the user's style requirements.
[0104] Step B2: Analyze the matching degree between the eigenvector of each aggregated environmental image and the target features to obtain the matching result.
[0105] Specifically, after determining the target features and the image synthesis strategy, it is necessary to evaluate the degree of fit between each aggregated environmental image and the target features. Since the eigenvectors of each original environmental image have been calculated before and the aggregated environmental images have been obtained through aggregation, these eigenvectors now become the basis for evaluation. Compare and analyze the eigenvector of each aggregated environmental image with the target features. For example, for the color feature, compare the similarity between the main color tone in the aggregated environmental image and the target color of the retro style; for the contrast feature, determine whether the contrast of the current aggregated environmental image is within the range required by the retro style. Through careful comparison of each feature dimension, comprehensively obtain the matching degree between each aggregated environmental image and the target features.
[0106] The analysis result of this matching degree can be presented in a quantitative manner to form the matching result. For example, the matching degree can be represented by a percentage. 90% indicates a high matching degree between the aggregated environmental image and the target features, while 30% indicates a low matching degree. This matching result provides a basis for subsequent selection of appropriate target environmental images for processing, helping the system determine which aggregated environmental images are more suitable for generating multimedia data that meets the user's style requirements through specific image synthesis strategies.
[0107] Step B3: Determine the target environmental image from the aggregated environmental images according to the matching result, and use the image synthesis strategy to process the target environmental image to generate the multimedia data corresponding to the style requirements.
[0108] Specifically, based on the obtained matching results, images with a relatively high degree of matching with the target features are selected from numerous aggregated environmental images as target environmental images. These target environmental images have a feature basis that better conforms to the user's style requirements and are more likely to generate multimedia data that meets the user's expectations after subsequent processing. The selected image synthesis strategy is used to process each feature dimension of the target environmental images according to the image synthesis strategy, and finally, multimedia data that meets the user's style requirements is generated. This multimedia data can be a panoramic view, where different target environmental images are stitched together according to a certain layout pattern to form a complete panoramic picture with a vintage style; it can also be a continuous video, where the target environmental images are processed sequentially and transitional effects are added to generate a continuous video with a vintage style, thereby providing a way for users to record and display travel images that meet their specific style requirements and solving the problem that picture processing in the prior art cannot meet the personalized style requirements of users.
[0109] In the embodiment of the present application, the target environmental images are processed using an image synthesis strategy to generate multimedia data corresponding to the style requirements, including the following steps B31 - B33:
[0110] Step B31, determine the corresponding data type according to the eigenvector of the target environmental image.
[0111] Specifically, the eigenvector is a quantitative representation of the features of the image in multiple dimensions. By analyzing it, the characteristics presented by the image data can be inferred, and then the corresponding data type can be determined. The data types are mainly divided into regular types and irregular types. For regular type data, the images have a relatively unified structure, color distribution, or geometric features, etc. For example, a series of photos taken from different angles of the same building may have certain regularities in composition, proportion, etc., making their data features show regularity. Irregular type data, on the contrary, has large differences in content, structure, etc. and there is no obvious unified rule. For example, randomly taken pictures of different scenes such as scenery, people, and food during a trip, their data features belong to the irregular type. By accurately judging the data type, it lays a foundation for subsequent selection of appropriate layout patterns and processing methods.
[0112] Step B32, obtain the layout pattern associated with the data type, and use the time arrangement rule to mark each target environmental image to obtain the target environmental image with a serial number identifier.
[0113] Specifically, after determining the data type, obtain the layout pattern associated with the data type according to existing knowledge or preset rules. The layout pattern determines the arrangement of the target environment images in the final multimedia data. If the data type is a regular type, the layout pattern can be a relatively neat and orderly arrangement, such as a specific row-column order or a symmetric structure. If it is an irregular type, the layout pattern can pay more attention to the visual relevance or narrativeness between images and combine the images in a more free and flexible way.
[0114] Meanwhile, to better organize and process the target environment images, use the time arrangement rule to mark each target environment image. This is because the pictures taken during the travel often have a time sequence, and this time sequence is very helpful for presenting the coherence and logic of the travel. Through the time arrangement rule, the system assigns a serial number identifier to each target environment image, and this serial number reflects the order of image shooting. For example, the picture taken earliest during the journey is marked as 1, and the subsequent ones are marked as 2, 3, etc. In this way, each target environment image becomes a target environment image with a serial number identifier, which not only contains its own image feature information but also has the position information in the time sequence, facilitating subsequent orderly processing according to the layout pattern.
[0115] Step B33: Process the target environment images with serial number identifiers according to the layout pattern to generate the multimedia data corresponding to the style requirements.
[0116] In the embodiment of the present application, step B33 includes: If the data type is a regular type, adjust the target environment images with serial number identifiers according to the layout parameters in the first layout pattern, and sequentially fill the adjusted target environment images into the specified positions of the first layout pattern according to the serial number identifiers to obtain the multimedia data; if the data type is an irregular type, analyze the positional relationship between each target environment image with a serial number identifier according to the second layout pattern, and fuse the target environment images based on the positional relationship and the serial number identifier to obtain the multimedia data.
[0117] Specifically, when the data type is determined to be a regular type, the target environmental images carrying serial number identifiers are processed according to the first layout mode. The first layout mode includes a series of layout parameters that determine the presentation of the images in the final multimedia data. For example, the layout parameters can specify that the size of the images is uniformly a certain fixed size or scaled according to a certain ratio; they can also specify the rotation angle of the images to ensure that all images have the same orientation when spliced. According to the layout parameters, each target environmental image carrying a serial number identifier is adjusted accordingly. After the adjustment of the target environmental images is completed, these images are filled into the specified positions of the first layout mode in sequence according to the serial number identifiers. For example, the first layout mode may specify a nine-grid layout. According to the serial numbers of the images, the image with serial number 1 is placed in the upper left corner of the nine-grid, the image with serial number 2 is placed on its right side, and so on, until all the target environmental images carrying serial number identifiers are accurately filled into the corresponding positions, finally forming multimedia data that conforms to the regular type layout.
[0118] This kind of multimedia data presents a neat and orderly visual effect, such as a panoramatic picture with regular splicing or a collection of pictures with specific arrangement rules, which is suitable for some style requirements with higher demands on structure and order, such as the style of a formal landscape picture album.
[0119] For non-regular type data, the second layout mode is adopted for processing. First, analyze the positional relationships among the target environmental images carrying serial number identifiers. This analysis is not only based on the serial numbers of the images but also comprehensively considers various factors such as the content, color, and visual focus of the images. For example, for a collection of images themed on travel experiences, which contain images of different scenes and different subjects, analyze the internal connections among these images, such as placing images of different details of the same scenic spot adjacent to each other, or combining images with echoing colors together to enhance visual coherence and logic. After analyzing the positional relationships, the target environmental images are fused based on these relationships and the serial number identifiers. During the fusion process, the boundaries of the images are blurred to make the transition between adjacent images more natural and avoid obvious splicing marks. For example, when fusing an image of a seaside landscape and an image of a person on the beach, the boundaries of the images are blurred to make the connection between the beach and the person look smoother. In addition, some transitional elements, such as lines and light and shadow effects, can be added as needed to further connect different images to generate a visually coordinated multimedia data that can reflect the travel characteristics and the user's style requirements.
[0120] This kind of multimedia data is more creative and narrative, such as forming a panoramatic picture with a unique narrative style or showing a continuous video of the travel process through a unique combination of images, meeting the user's style requirements for personalization and creativity.
[0121] In this embodiment, a vehicle environment image acquisition and processing system is also provided. As Figure 4 shown, the system includes: a user side and a vehicle side. Among them, the user side is connected to the trigger acquisition module in the vehicle side;
[0122] The user side is used to send relevant instructions and set parameters to the trigger acquisition module of the vehicle side to control the image acquisition process of the vehicle side;
[0123] The vehicle side is used to receive the relevant instructions and set parameters sent by the user side, and perform image acquisition, processing, and synthesis operations according to the relevant instructions and set parameters to generate multimedia data that meets the user's needs.
[0124] In the embodiment of the present application, the vehicle side includes a trigger acquisition module, an image acquisition algorithm, an image feature model, an image acquisition module, and an image synthesis module;
[0125] The trigger acquisition module is used to monitor the change of external environment parameters according to the trigger parameters set by the user. When the preset conditions are met, it controls the cameras (such as camera 1, camera 2, camera 3) associated with the trigger acquisition module to perform acquisition operations, obtains a plurality of original environment images, and sends the original environment images to the image acquisition algorithm;
[0126] The feature extraction module is used to extract features from a plurality of original environment images, obtain the feature value vectors of each original environment image, and send the feature value vectors to the image feature model;
[0127] The image feature model is used to receive the received feature value vectors sent by the image acquisition algorithm and store the feature value vectors of each original environment image in the image feature model;
[0128] The image aggregation module is used to obtain the feature value vectors of the original environment images in the image feature model, and construct a plurality of aggregated image groups from the original environment images with similar features by means of similarity calculation, initial grouping, and cyclic assignment and group construction, and send the aggregated image groups to the image synthesis module;
[0129] The image synthesis module is used to receive the aggregated image groups sent by the image aggregation module, and generate regular or irregular multimedia data that meets the user's style requirements based on the aggregated image groups through style requirement and target feature determination, feature matching and target image screening, and data type judgment and corresponding processing.
[0130] Specifically, as Figure 5As shown in the figure, the workflow of the vehicle environment image acquisition and processing system includes: First, the client sends relevant instructions and setting parameters to the trigger acquisition module on the vehicle side. These parameters cover various environmental factors such as light, color temperature, and brightness. Users can make personalized settings according to their own needs to control the image acquisition process on the vehicle side.
[0131] Second, after receiving the instructions and parameters sent by the client, the trigger acquisition module on the vehicle side monitors the changes in external environmental parameters in real time. When the external environmental parameters meet the preset conditions set by the user, the trigger acquisition module will control the associated cameras (such as Camera 1, Camera 2, Camera 3) to perform the acquisition operation, thereby obtaining multiple original environmental images and sending the original environmental image to the image acquisition algorithm.
[0132] Third, after receiving the original environmental image, the image acquisition algorithm's feature extraction module filters out the effective environmental images from the original environmental images, extracts the image features under different indicators of the effective environmental images, constructs the feature value vector of the original environmental image, and sends the feature value vector of each original environmental image to the image feature model.
[0133] The image feature model stores the feature value vector of each original environmental image. At the same time, the image aggregation module obtains the feature value vectors of the original environmental images in the image feature model, and constructs multiple aggregated image groups of the original environmental images with similar features through similarity calculation, initial grouping, and cyclic assignment and group construction.
[0134] Finally, after receiving the aggregated image groups sent by the image aggregation module, the image synthesis module determines the target features and the corresponding image synthesis strategy according to the style requirements input by the user. Analyze the matching degree between the feature value vector of each aggregated environmental image and the target features, obtain the matching result through quantization, and determine the target environmental image from the aggregated environmental images according to the matching result. Then, determine its corresponding data type according to the feature value vector of the target environmental image, which is divided into regular type and irregular type. If the data type is regular type, generate regularly arranged multimedia data according to the first layout mode; if the data type is irregular type, fuse the target environmental image according to the second layout mode to generate irregular multimedia data that meets the user's style requirements, providing users with diverse multimedia display effects and meeting the user's personalized recording and sharing needs for environmental images during vehicle driving.
[0135] In this embodiment, a vehicle multimedia data synthesis device is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0136] This embodiment provides a vehicle multimedia data synthesis device, as Figure 6 shown, including:
[0137] An acquisition module 61, configured to acquire a plurality of original environment images;
[0138] A construction module 62, configured to construct a plurality of aggregated image groups by using the original environment images, where the plurality of original environment images in each aggregated image group are similar environment images;
[0139] An aggregation module 63, configured to aggregate the original environment images in each aggregated image group to obtain an aggregated environment image corresponding to each aggregated image group;
[0140] A processing module 64, configured to obtain the style requirements input by the user, and process the aggregated environment image by using an image synthesis strategy corresponding to the style requirements to generate multimedia data corresponding to the style requirements.
[0141] In an alternative embodiment of the present application, the construction module 62 is configured to calculate the eigenvalue vector of each original environment image; determine the environment images with similar features among the plurality of original environment images based on the eigenvalue vector, and construct a plurality of aggregated image groups by using the environment images with similar features.
[0142] In an alternative embodiment of the present application, the construction module 62 includes an extraction sub-module and an assignment sub-module;
[0143] The extraction sub-module is configured to extract the pixel information in the original environment image; determine whether the original environment image is a valid image according to the pixel information; when the original environment image is a valid image, extract the image features under different metrics in the original environment image; and construct an eigenvalue vector by using the image features under different metrics in the original environment image.
[0144] An allocation sub-module is configured to select a first environmental image from the original environmental images, and construct a first aggregated image group based on the first environmental image; use the eigenvalue vector to select a second environmental image similar to the first environmental image from multiple original environmental images, and allocate the second environmental image to the first aggregated image group; select a third environmental image from the remaining environmental images, and construct a second aggregated image group based on the third environmental image, where the remaining environmental images are the environmental images in the original environmental images except the first environmental image and the second environmental image; use the eigenvalue vector to select a fourth environmental image similar to the third environmental image from multiple remaining environmental images, and allocate the fourth environmental image to the second aggregated image group; repeatedly execute the image allocation operation until all the original environmental images are allocated, and obtain multiple aggregated image groups.
[0145] The allocation sub-module is further configured to calculate the similarity between the eigenvalue vector of the first environmental image and the eigenvalue vectors of other environmental images in the original environmental images except the first environmental image, to obtain a similarity matrix; and determine at least one second environmental image similar to the first environmental image from the other environmental images according to the similarity matrix.
[0146] In an optional embodiment of the present application, a processing module 64 is configured to determine a target feature and an image synthesis strategy according to the style requirement; analyze the matching degree between the eigenvalue vector of each aggregated environmental image and the target feature, to obtain a matching result; determine a target environmental image from the aggregated environmental images according to the matching result, and process the target environmental image by using the image synthesis strategy to generate multimedia data corresponding to the style requirement.
[0147] In an optional embodiment of the present application, the processing module 64 is configured to determine the corresponding data type according to the eigenvalue vector of the target environmental image; obtain the layout mode associated with the data type, and use the time arrangement rule to mark each target environmental image, to obtain a target environmental image carrying a serial number identifier; process the target environmental image carrying the serial number identifier according to the layout mode to generate multimedia data corresponding to the style requirement.
[0148] In an optional embodiment of the present application, the processing module 64 is configured to, if the data type is a regular type, adjust the target environmental image carrying the serial number identifier according to the layout parameters in the first layout mode, and sequentially fill the adjusted target environmental images into the specified positions of the first layout mode according to the serial number identifier to obtain multimedia data; if the data type is an irregular type, analyze the positional relationship between each target environmental image carrying the serial number identifier according to the second layout mode, and fuse the target environmental images based on the positional relationship and the serial number identifier to obtain multimedia data.
[0149] Please refer to Figure 7 , Figure 7It is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 7 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system).
[0150] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0151] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0152] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a kind of landing page of a small program, etc. In addition, the memory 20 can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0153] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.
[0154] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.
[0155] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0156] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for synthesizing vehicle multimedia data, characterized in that: The method comprises: Acquire multiple original environment images; Constructing a plurality of aggregated image groups using the original environment image, wherein the plurality of original environment images in each of the aggregated image groups are similar environment images; Aggregating the original environment images in each of the aggregated image groups to obtain an aggregated environment image corresponding to each aggregated image group; The style requirement input by the user is obtained, and the aggregated environment image is processed using an image synthesis strategy corresponding to the style requirement to generate multimedia data corresponding to the style requirement.
2. The method according to claim 1, characterized in that The step of constructing a plurality of aggregated image groups by using the original environment image comprises: Calculate the eigenvalue vector of each of the original environment images; Based on the eigenvalue vector, environmental images with similar features are determined from a plurality of original environmental images, and a plurality of aggregated image groups are constructed using the environmental images with similar features.
3. The method according to claim 2, characterized in that The step of calculating the eigenvalue vector of each original environment image comprises: Extracting pixel information from the original environment image; Determining whether the original environment image is a valid image according to the pixel information; When the original environment image is a valid image, extracting image features under different indicators in the original environment image; The eigenvalue vector is constructed by using image features under different indicators in the original environment image.
4. The method according to claim 2, characterized in that: The step of determining environmental images with similar features from a plurality of original environmental images based on the eigenvalue vectors, and constructing a plurality of aggregated image groups using the environmental images with similar features, comprises: Selecting a first environment image from the original environment image, and constructing a first aggregated image group based on the first environment image; selecting a second environment image similar to the first environment image from the plurality of original environment images by using the eigenvalue vector, and assigning the second environment image to the first aggregated image group; Selecting a third environment image from the remaining environment images, and constructing a second aggregated image group based on the third environment image, wherein the remaining environment images are environment images other than the first environment image and the second environment image in the original environment images; selecting a fourth environment image similar to the third environment image from the plurality of remaining environment images by using the eigenvalue vector, and assigning the fourth environment image to the second aggregated image group; The image allocation operation is repeatedly performed until all the original environment images are allocated, thereby obtaining a plurality of aggregated image groups.
5. The method according to claim 4, characterized in that The selecting a second environment image similar to the first environment image from the plurality of original environment images by using the eigenvalue vector comprises: Calculating the similarity between the eigenvalue vector of the first environment image and the eigenvalue vectors of other environment images in the original environment image except the first environment image to obtain a similarity matrix; At least one second environment image similar to the first environment image is determined from the other environment images according to the similarity matrix.
6. The method according to claim 1, characterized in that The step of processing the aggregated environment image by using the image synthesis strategy corresponding to the style requirement to generate multimedia data corresponding to the style requirement includes: Determining target features and image synthesis strategies according to the style requirements; Analyzing the matching degree between the feature value vector of each of the aggregated environment images and the target feature to obtain a matching result; A target environment image is determined in the aggregated environment image according to the matching result, and the target environment image is processed using the image synthesis strategy to generate multimedia data corresponding to the style requirement.
7. The method according to claim 6, characterized in that The step of processing the target environment image by using the image synthesis strategy to generate multimedia data corresponding to the style requirement includes: Determining a corresponding data type according to the eigenvalue vector of the target environment image; Acquire the layout mode associated with the data type, and mark each of the target environment images using a time arrangement rule to obtain a target environment image carrying a serial number identifier; The target environment image carrying the serial number identifier is processed according to the layout mode to generate multimedia data corresponding to the style requirement.
8. The method according to claim 7, characterized in that The step of processing the target environment image carrying the serial number identifier according to the layout mode to generate multimedia data corresponding to the style requirement includes: If the data type is a rule type, the target environment image carrying the serial number identifier is adjusted according to the layout parameters in the first layout mode, and the adjusted target environment image is sequentially filled into the designated position of the first layout mode according to the serial number identifier to obtain multimedia data; If the data type is a non-regular type, the positional relationship between the target environment images carrying serial number identifiers is analyzed according to the second layout mode, and the target environment images are fused based on the positional relationship and the serial number identifiers to obtain multimedia data.
9. A synthesis device for vehicle multimedia data, characterized in that: The device comprises: An acquisition module, used for acquiring multiple original environment images; A construction module, configured to construct a plurality of aggregated image groups using the original environment image, wherein the plurality of original environment images in each of the aggregated image groups are similar environment images; An aggregation module, used for aggregating the original environment images in each of the aggregated image groups to obtain an aggregated environment image corresponding to each aggregated image group; The processing module is used to obtain the style requirement input by the user, and use the image synthesis strategy corresponding to the style requirement to process the aggregated environment image to generate multimedia data corresponding to the style requirement.
10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 8 by executing the computer instructions.