Multimedia resource storage method and device, equipment, storage medium and program product
By segmenting image content and generating descriptive text using image recognition technology, the problem of large storage space consumption of multimedia resources is solved, achieving storage space saving and convenient image editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2024-06-24
- Publication Date
- 2026-04-21
AI Technical Summary
As electronic devices enhance their camera capabilities and multimedia resources achieve higher resolution, they also require increasingly larger storage spaces, making it difficult for existing technologies to effectively save storage space.
Image content is divided into a first element and a second element using image recognition technology. Image description text for the second element is generated and stored along with the pixel information of the first element, thus reducing storage space usage.
It effectively saves storage space for multimedia resources while supporting image editing and personalization needs, thus improving storage space utilization efficiency.
Smart Images

Figure CN118827896B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data storage technology, specifically relating to a multimedia resource storage method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the continuous development of intelligent electronic devices, more and more users are using electronic devices to capture images or videos and other multimedia resources to record their lives. As the camera capabilities of electronic devices continue to improve, the clarity of multimedia resources is getting higher and higher, and correspondingly, the storage space they occupy is also getting larger and larger. Summary of the Invention
[0003] The purpose of this application is to provide a multimedia resource storage method, apparatus, device, storage medium, and program product that can save storage space occupied by multimedia resources.
[0004] In a first aspect, embodiments of this application provide a multimedia resource storage method, the method comprising:
[0005] Display multimedia resources, which include at least one frame of image;
[0006] For each frame of the first image in the multimedia resource, the image content of the first image is divided to obtain a first element and a second element. The second element includes at least a portion of the image content of the first image, and the first element includes the image content of the first image excluding the second element.
[0007] Generate image description text based on the image content of the second element;
[0008] For each frame of the first image in the multimedia resource, store the image description text and the pixel information corresponding to the image content of the first element.
[0009] Secondly, embodiments of this application provide a multimedia resource storage device, the device comprising:
[0010] The first display module is used to display multimedia resources, which include at least one frame of image.
[0011] The first determining module is used to divide the image content of each frame of the first image in the multimedia resource to obtain a first element and a second element. The second element includes at least a portion of the image content of the first image, and the first element includes the image content of the first image excluding the second element.
[0012] The first generation module is used to generate image description text based on the image content of the second element;
[0013] The first storage module is used to store the image description text and the pixel information corresponding to the image content of the first element for each frame of the first image in the multimedia resource.
[0014] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions implementing the steps of the method as described in the first aspect when executed by the processor.
[0015] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, they implement the steps of the method as described in the first aspect.
[0016] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method as described in the first aspect.
[0017] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.
[0018] In this embodiment, multimedia resources can be displayed; at least a portion of the image content of each frame of the first image in the multimedia resources is determined as a second element, and the image content of the first image excluding the second element is determined as a first element; image description text is generated based on the image content of the second element; and pixel information corresponding to the image description text and the image content of the first element is stored. In this way, all or part of the image content of the first image can be stored using image description text, while the remaining image content stores its pixel information. Compared to storing the pixel information of the entire first image, this method occupies less storage space, thereby saving storage space occupied by multimedia resources. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the multimedia resource storage method provided in an embodiment of this application;
[0020] Figure 2 This is one of the interface diagrams of the multimedia resource storage method provided in the embodiments of this application;
[0021] Figure 3 This is a second schematic diagram of the interface of the multimedia resource storage method provided in the embodiments of this application;
[0022] Figure 4 This is a schematic diagram illustrating the feature changes in the multimedia resource storage method provided in the embodiments of this application;
[0023] Figure 5 It is a schematic structural diagram of a multimedia resource storage device provided by an embodiment of the present application;
[0024] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0025] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0026] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0027] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character "and / or" generally indicates that the associated objects are in an "or" relationship.
[0028] Next, in conjunction with the accompanying drawings, the multimedia resource storage method provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0029] Figure 1 It is a schematic flowchart of a multimedia resource storage method provided by an embodiment of the present application. The multimedia resource storage method may include:
[0030] Step 101, the electronic device displays a multimedia resource, and the multimedia resource includes at least one frame of image.
[0031] In step 101, the multimedia resource may be an image or a video. When the electronic device displays the multimedia resource, it may be one or more preview images displayed when the electronic device captures an image or a video. It may also be that the user selects an image or a video from the album of the electronic device for display. Hereinafter, the multimedia resource is taken as an image for illustration.
[0032] Step 102: For each frame of the first image in the multimedia resource, divide the image content of the first image to obtain a first element and a second element. The second element includes at least a portion of the image content of the first image, and the first element includes the image content of the first image excluding the second element.
[0033] In step 102, when the electronic device displays the first image, multiple elements contained in the image content of the first image can be identified by image recognition technology, and these elements can be classified into the first element and the second element.
[0034] In some examples, the categories of the first and second elements can be predefined. For instance, the second element could include a solid-color or simple background image such as a sky, clouds, lake, grass, trees, or street, while the first element could include the main subject of the image, such as people, animals, or buildings. Multiple elements identified in the first image can be categorized into the first and second elements based on their classifications and the categories of the first and second elements. For example... Figure 2 As shown, the image content of the first image 201 may include people, animals, trees, blue sky, and white clouds. Therefore, the image content corresponding to trees, blue sky, and white clouds can be determined as the image content of the second element 202, and the image content corresponding to people and animals can be determined as the image content of the first element 203.
[0035] In other examples, when displaying a first image, the system may receive a selection of the image content within that first image, thereby determining either a first element or a second element. It is understood that the user can identify the entire image content of the first image as the second element; in other words, there is a possibility that the first image contains only the second element.
[0036] In the embodiments of this application, the specific division method of the first element and the second element can be set according to actual needs, and is not specifically limited here.
[0037] Step 103: Generate image description text based on the image content of the second element.
[0038] Understandably, in one application scenario, a user is playing outdoors with their family, and during this time, the user takes photos of their family members. At this point, the electronic device's camera preview interface displays something like this. Figure 2The first image 201 shown includes elements such as people (i.e., family members), animals (i.e., pets), trees, blue sky, and white clouds. The key information to be displayed in the first image 201 is the people and animals, equivalent to the first element 203 mentioned above. Since the image content of elements such as trees, blue sky, and white clouds displays almost identically across different images, they are usually not the key information to be displayed, equivalent to the second element 201 mentioned above. Therefore, when storing the first image 201, this part of the image content can be optimized to save storage space.
[0039] Based on this, in step 103, AI Generated Content (AIGC) technology can be used to generate textual descriptions of the image content of the second element, thus creating image description text. It is understood that the image description text may include the position coordinates of the second element in the first image, as well as textual information describing the image content itself.
[0040] For example, such as Figure 2 As shown, the image description text may include location coordinates, such as blue sky and white clouds at the top of the first image 201, and may also include text information, such as "a white cloud is floating in the blue sky".
[0041] Step 104: For each frame of the first image in the multimedia resource, store the image description text and the pixel information corresponding to the image content of the first element.
[0042] In step 104, when storing the first image, image description text can be used to replace the pixel information corresponding to the image content of the second element. In other words, both the image description text and the pixel information corresponding to the image content of the first element can be stored. That is, the image data of the first image can be saved in the form of "pixel information + image description text". The pixel information can be pixel values or RGB values; no specific limitation is made here.
[0043] Understandably, for videos, each frame can be saved using the above method, in the form of "pixel information + image description text", thereby saving overall video storage space.
[0044] It is also understandable that, such as Figure 2 As shown, in scenarios where multimedia resources are displayed on the shooting preview interface, the first image 201 can be processed and saved using the above method when the user clicks the first control 204, such as the "AI Save" control displayed on the interface. Figure 3As shown, in scenarios where multimedia resources are images or videos previously saved in the photo album of an electronic device, when the user clicks the second control 302, such as the "AI Storage Optimization" control displayed on the interface, the multimedia resource 301 selected by the user can be processed and stored in the manner described above. Then, the data of the multimedia resource 301 that was originally stored can be deleted to save storage space.
[0045] In this embodiment, the multimedia resource storage method can display multimedia resources; determine at least a portion of the image content of any frame of a first image in the multimedia resources as a second element, and determine the image content of the first image excluding the second element as a first element; generate image description text based on the image content of the second element; and store the pixel information corresponding to the image description text and the image content of the first element. In this way, all or part of the image content of the first image can be stored using image description text, while the remaining image content stores its pixel information. Compared to storing the pixel information of the entire first image, this method occupies less storage space, thereby saving storage space occupied by multimedia resources.
[0046] In some embodiments, step 103 above may include the following steps:
[0047] Generate and display the initial image description text based on the image content of the second element;
[0048] Receive edit input for the initial description text of the image;
[0049] In response to edit input, the image description text is obtained.
[0050] In this embodiment, an initial image description text can be generated based on the image content of the second element. The initial image description text can be generated by using AIGC technology to describe the image content of the second element in the first image.
[0051] Understandably, during the actual shooting process, users may have dissatisfaction with the image content and need to make modifications. For example, a user might feel that a person's expression is not good enough; for instance, the expression in the first image might be a bit "gloomy," and the user might want to change it to a "smile." Or, a user might feel that there are not enough clouds in the first image and want to increase the number of clouds from one to two. Alternatively, a user might want to adjust the position of a certain element in the first image. Therefore, if a user wants to modify the state (expression, posture, shape, etc.) of elements in an image, or the layout of elements, or the mood (joyful, peaceful, desolate, etc.) of elements, they can do so by receiving editing input for the initial image description text, modifying the initial image description text, and obtaining the modified image description text, thus achieving the purpose of modifying the image content.
[0052] For example, if the entire content of the first image is used as the second element, the initial image description text could be: "In the park, there is a pretty little girl, but her expression is a bit melancholy. Next to her sits her pet dog. The scenery in the park is beautiful, with large trees and a white cloud floating in the blue sky." When this initial image description text is displayed, editing input can be received. Responding to the editing input, the image description text is obtained, and the edited image description text could be: "In the park, there is a pretty little girl with a smiling expression. Next to her sits her pet dog. The scenery in the park is beautiful, with large trees and a white cloud floating in the blue sky." Alternatively, only the part that needs modification can be used as the second element. For example, if the person is determined as the second element, the initial image description text could be: "This little girl is pretty, but her expression is a bit melancholy." Responding to editing input, the edited image description text could be: "This little girl is pretty and has a smiling expression." Later, when restoring the image based on the image description text, an image including the smiling girl can be obtained.
[0053] In this way, AIGC technology can be used to describe the image content with text, resulting in the initial image description text. By editing this initial image description text, the image content can be modified, thus enabling image editing based on the image description text and improving the diversity and convenience of image editing.
[0054] In some embodiments, after step 104 above, the multimedia resource storage method may further include:
[0055] Receive the first input;
[0056] In response to the first input, the element content of the second element is generated based on the image description text;
[0057] Based on the pixel information, the image content of the first element is obtained through parsing.
[0058] Based on the element content of the second element and the image content of the first element, a second image is generated, and the image content of the second image corresponds to the image content of the first image;
[0059] The second image is displayed.
[0060] In this embodiment, a first input from a user to the electronic device can be received, which may indicate that the user wants to view multimedia resources stored in the manner described above.
[0061] In response to the first input, the electronic device can parse the stored data of the multimedia resource. Taking the first image as an example, when it is detected that the stored data of the first image includes two parts: image description text and pixel information corresponding to the image content of the first element, the element content of the second element can be generated based on the image description text.
[0062] For example, such as Figure 2 As shown, the blue sky and white cloud elements in the first image 201 are stored in the form of image description text "a white cloud is floating in the blue sky". When reading and displaying the first image 201, it is necessary to generate the corresponding image content, i.e. the element content of the second element, again through AIGC technology based on the image description text.
[0063] Understandably, the phrase "a white cloud floating in a blue sky" includes the elements "sky" and "cloud," with the adjective "blue" describing the sky as very blue, a deep blue, and the quantifier "a cloud" and the adjective "white" describing the cloud. Once the electronic device acquires this information—including the element objects, adjectives, and quantifiers—it can use AIGC technology to generate an effect similar to a previously captured photo of a blue sky and white clouds. It's also understandable that the content of the second element can be associated with its display location coordinates.
[0064] For the pixel information, the pixels can be directly parsed and rendered to obtain the image content of the first element for display. A second image can be generated and displayed based on the element content of the second element and the image content of the first element, with the image content of the second image corresponding to the image content of the first image. For example, the element content of the second element and the image content of the first element can be merged and then displayed.
[0065] The displayed image is a second image corresponding to the first image. It can be understood that because the image content of the second element in the second image is generated by reverse engineering the image description text generated from the image content of the second element in the first image, the image content of the second element in the second image is close to the image content of the second element in the first image. Furthermore, the image content of the first element in the second image is exactly the same as the image content of the first element in the first image; therefore, the overall image content of the second image is close to the overall image content of the first image. In other words, the second image can be considered either the first image or an image edited based on the first image.
[0066] In this way, users can view multimedia resources stored in the form of "pixel information + image description text" at any time, saving storage space for multimedia resources without affecting their need to record life through storing multimedia resources.
[0067] In some embodiments, step 102 above may include the following steps:
[0068] While displaying the first image, receive a second input to the first image;
[0069] In response to the second input, a first element in the first image associated with the second input is determined;
[0070] The image content other than the first element in the first image is determined as the second element.
[0071] In this embodiment, while displaying the first image, a second input to the first image can be received. This second input can be a user's click on the first image, a voice command input by the user, or a specific gesture input by the user. The specific gesture can be determined according to actual usage needs, and this embodiment does not limit this. The specific gesture in this embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-tap gesture. The click input in this embodiment can be a single-click input, a double-tap input, or any number of clicks, and can also be a long-press input or a short-press input.
[0072] In response to a second input, a first element in the first image associated with the second input can be determined. For example, such as Figure 2 As shown, the user can long-press the area of the person in the first image 201 to designate the person as the first element 203. In other words, by using the person as key information in the first image 201, this part will directly store the corresponding pixel information when storing the first image 201 later.
[0073] The image content in the first image other than the first element is determined as the second element. For example, all other parts not selected by the user are treated as the second element 202, and these parts are subsequently stored using image description text described by AIGC technology.
[0074] This provides users with the option to operate as they see fit, saving multimedia resource storage space while meeting users' personalized needs.
[0075] In some embodiments, the multimedia resource includes N frames of images, where N is an integer greater than 1. After step 101 above, the multimedia resource storage method may further include:
[0076] Determine the target frame image from the multimedia resources, and M frames of images with the target frame image as the starting frame image, where M is less than or equal to N;
[0077] Determine the third element from the target frame image;
[0078] Based on the image content of the third element in the M-frame image, identify the feature changes of the third element, wherein the feature changes include at least one of position change, pose change, or state change;
[0079] Based on the feature changes of the third element, generate a feature change description text;
[0080] Store the target image frame and the text describing the feature changes.
[0081] In this embodiment, when the multimedia resource is video, the differences between multiple frames can also be described using AIGC technology, further reducing the video storage space.
[0082] The target frame image and the M-frame images, which use the target frame image as the starting frame image, can be determined from multimedia resources. For example, for a video recorded through a shooting preview interface, the starting image of the shooting preview interface can be considered a target frame image. For videos previously saved in the album, the target frame image can be determined based on the user's selection. The target frame image can be used as the starting frame image to obtain multiple frames, including the target frame image.
[0083] Taking the target frame image as the starting image of the shooting preview interface as an example, during the video recording process, compared with the target frame image, the elements in the subsequent images may change dynamically over time. For example, the position, state, or posture of the elements between multiple frames may change. Therefore, electronic devices can use AIGC technology to analyze the type and characteristics of the elements, predict or analyze the rules and characteristics of the changes in element features, and use AIGC text language to replace description.
[0084] The third element can be determined from the target frame image. It can be understood that the third element can be the second element mentioned above, the first element mentioned above, or a part of the second or first element; no specific limitation is made here. It can also be understood that the third element is often included in the M-frame image that uses the target frame image as the starting frame image.
[0085] For example, we can analyze the elements contained in the target frame image, such as cars, roads, and buildings. Based on these elements, we can predict the intent of the video scene involved, such as "a car is driving on the road," in which case "car" can be identified as the third element. Alternatively, we can determine the third element in the target image frame based on the user's selection.
[0086] Based on the image content of the third element in the M-frame image, feature changes of the third element can be identified, and feature change description text can be generated based on these feature changes. Feature changes can include at least one of position changes, pose changes, or state changes.
[0087] For example, when a red car is driving on the road, recording a video will generate multiple frames. We can analyze the changes in the position of this element in different images at different times. By roughly converting distance and time, we can obtain the element's motion, such as "driving at a constant speed." The car's position coordinates will then change linearly over time. Therefore, when saving the "car" element, we can dynamically calculate its coordinates and describe it using AIGC technology. The generated feature change description text could be something like, "A red car is currently driving at a constant speed of 80 km / h on the highway."
[0088] For example, if the third element in the target frame image is "Person A", then the adjustment time range for "Person A" can be determined to be "the video within the next 10 seconds". In other words, the M-frame images are all images starting from the target frame image, whose timestamps are within 10 seconds after the timestamp of the target frame image, and "Person A" in these images will be adjusted. If the feature change of "person A" is identified as the addition of an action, such as "jumping" or "speaking," these actions are carried by body parts; for example, "jumping" corresponds to the "legs," and "speaking" corresponds to the "mouth." The local position of "person A" can then be adjusted. Action postures can be constructed; for example, "jumping" and "speaking" have corresponding action model trajectories, such as up / down, forward movement, and frequency. These action model trajectories are then applied to "person A." Based on the trajectory position and frequency, these are applied to images at different times, resulting in corresponding feature changes in "person A" across multiple frames. Using AIGC technology, these changes can be described, and the generated feature change description text could be, "Based on the target frame image, in the following 10 seconds of video, person A added a 'jumping' or 'speaking' action."
[0089] For example, from the image content of the third element in an M-frame image, the characteristic change of the third element can be identified as a continuous deepening of its color. For instance... Figure 4 As shown, the feature change description text generated by AIGC technology can be "In the four frames starting from the first frame, the color of element 1 continuously deepens".
[0090] It can store target image frames and feature change description text. Understandably, the target image frames can be stored using the same method as the first image mentioned above, or all pixel information of the target frame image can be directly saved; the choice depends on actual needs and is not specifically limited here. The stored target image frames and feature change description text can be used as storage data for M frames, with the target frame image as the starting frame image.
[0091] To view a video segment including the M-frame image later, the pixel information of the rendered target frame image can be directly parsed to obtain the image content of the target frame image. Then, based on the feature change description text, the feature changes of the third element are obtained using AIGC technology. The content of the third element in the image content of the target frame image is replaced with the content of the third element based on the feature changes of the third element, thus obtaining the image content of the M-1 frame image.
[0092] For example, the image content of the target frame image may include a road, trees beside the road, and a red car. The descriptive text of the feature change could be "A red car is currently traveling at a constant speed of 80 km / h on the road." At this point, the displacement of the red car in adjacent frames can be calculated based on the time interval and speed between adjacent frames. Thus, while keeping the road and trees beside the road unchanged, only the position of the red car in each frame can be changed, ultimately resulting in the image content of M frames, including the target frame image, which can then be displayed to achieve the purpose of playing a video segment.
[0093] In this way, compared to storing all pixel information of the M-frame images, storing the target image frame and the text describing the feature changes can effectively reduce the storage space occupied by the M-frame images, and further achieve the goal of saving multimedia resource storage space.
[0094] The multimedia resource storage method provided in this application can be executed by a multimedia resource storage device. This application uses an example of a multimedia resource storage device executing the multimedia resource storage method to illustrate the multimedia resource storage device provided in this application.
[0095] like Figure 5 As shown, the multimedia resource storage device 500 provided in this embodiment may include:
[0096] The first display module 501 is used to display multimedia resources, which include at least one frame of image.
[0097] The first determining module 502 is used to divide the image content of each frame of the first image in the multimedia resource to obtain a first element and a second element. The second element includes at least a part of the image content of the first image, and the first element includes the image content of the first image other than the second element.
[0098] The first generation module 503 is used to generate image description text based on the image content of the second element;
[0099] The first storage module 504 is used to store the image description text and the pixel information corresponding to the image content of the first element for each frame of the first image in the multimedia resource.
[0100] In this way, all or part of the image content in the first image can be stored through image description text, while the remaining image content stores its pixel information. This takes up less storage space than storing the pixel information of the entire first image, thereby saving storage space occupied by multimedia resources.
[0101] In some embodiments, the first generation module 503 can also be used for:
[0102] Generate and display the initial image description text based on the image content of the second element;
[0103] Receive edit input for the initial description text of the image;
[0104] In response to edit input, the image description text is obtained.
[0105] In this way, AIGC technology can be used to describe the image content with text, resulting in an initial image description text. By editing this initial image description text, the image content can be modified, thus improving the diversity and convenience of image editing.
[0106] In some embodiments, the multimedia resource storage device 500 may further include:
[0107] The first receiving module is used to receive the first input;
[0108] The second generation module is used to generate the element content of the second element in response to the first input, based on the image description text.
[0109] The parsing module is used to parse the image content of the first element based on the pixel information;
[0110] The third generation module is used to generate a second image based on the element content of the second element and the image content of the first element, wherein the image content of the second image corresponds to the image content of the first image;
[0111] The second display module is used to display the second image.
[0112] In this way, users can view multimedia resources stored in the form of "pixel information + image description text" at any time, saving storage space for multimedia resources without affecting their need to record life through storing multimedia resources.
[0113] In some embodiments, the first determining module 502 can also be used for:
[0114] While displaying the first image, receive a second input to the first image;
[0115] In response to the second input, a first element in the first image associated with the second input is determined;
[0116] The image content other than the first element in the first image is determined as the second element.
[0117] This provides users with the option to operate as they see fit, saving multimedia resource storage space while meeting users' personalized needs.
[0118] In some embodiments, the multimedia resource includes N frames of images, where N is an integer greater than 1, and the multimedia resource storage device 500 may further include:
[0119] The second determining module is used to determine the target frame image from the multimedia resources, and M frame images with the target frame image as the starting frame image, where M is less than or equal to N;
[0120] The third determining module is used to determine the third element from the target frame image;
[0121] The recognition module is used to recognize feature changes of the third element based on the image content of the third element in the M-frame image. The feature changes include at least one of position change, pose change, or state change.
[0122] The fourth generation module is used to generate feature change description text based on the feature changes of the third element;
[0123] The second storage module is used to store the target image frame and the text describing the feature changes.
[0124] In this way, compared to storing all pixel information of the M-frame images, storing the target image frame and the text describing the feature changes can effectively reduce the storage space occupied by the M-frame images, and further achieve the goal of saving multimedia resource storage space.
[0125] The multimedia resource storage device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0126] The multimedia resource storage device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0127] The multimedia resource storage device provided in this application embodiment can achieve... Figures 1 to 4 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0128] Optionally, such as Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601 and a memory 602. The memory 602 stores a program or instructions that can run on the processor 601. When the program or instructions are executed by the processor 601, they implement the various steps of the above-described multimedia resource storage method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0129] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.
[0130] Figure 7 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0131] The electronic device 700 includes, but is not limited to, components such as: radio frequency unit 701, network module 702, audio output unit 703, input unit 704, sensor 705, display unit 706, user input unit 707, interface unit 708, memory 709, and processor 710.
[0132] Those skilled in the art will understand that the electronic device 700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0133] The display unit 706 can be used to display multimedia resources, which include at least one frame of image.
[0134] The processor 710 can be used for:
[0135] For each frame of the first image in the multimedia resource, the image content of the first image is divided to obtain a first element and a second element. The second element includes at least a portion of the image content of the first image, and the first element includes the image content of the first image excluding the second element.
[0136] Generate image description text based on the image content of the second element;
[0137] For each frame of the first image in the multimedia resource, store the image description text and the pixel information corresponding to the image content of the first element.
[0138] In this way, all or part of the image content in the first image can be stored through image description text, while the remaining image content stores its pixel information. This takes up less storage space than storing the pixel information of the entire first image, thereby saving storage space occupied by multimedia resources.
[0139] In some embodiments, the processor 710 can also be used for:
[0140] Generate and display the initial image description text based on the image content of the second element;
[0141] Receive edit input for the initial description text of the image;
[0142] In response to edit input, the image description text is obtained.
[0143] In this way, AIGC technology can be used to describe the image content with text, resulting in an initial image description text. By editing this initial image description text, the image content can be modified, thus improving the diversity and convenience of image editing.
[0144] In some embodiments, the user input unit 707 may be used to: receive a first input;
[0145] The processor 710 can also be used for:
[0146] In response to the first input, the element content of the second element is generated based on the image description text;
[0147] Based on the pixel information, the image content of the first element is obtained through parsing.
[0148] Based on the element content of the second element and the image content of the first element, a second image is generated, and the image content of the second image corresponds to the image content of the first image;
[0149] The display unit 706 can be used to display a second image.
[0150] In this way, users can view multimedia resources stored in the form of "pixel information + image description text" at any time, saving storage space for multimedia resources without affecting their need to record life through storing multimedia resources.
[0151] In some embodiments, the user input unit 707 may be used to: receive a second input to the first image when the first image is displayed;
[0152] The processor 710 can also be used for:
[0153] In response to the second input, a first element in the first image associated with the second input is determined;
[0154] The image content other than the first element in the first image is determined as the second element.
[0155] This provides users with the option to operate as they see fit, saving multimedia resource storage space while meeting users' personalized needs.
[0156] In some embodiments, the multimedia resource includes N frames of images, where N is an integer greater than 1.
[0157] The processor 710 can also be used for:
[0158] Determine the target frame image from the multimedia resources, and M frames of images with the target frame image as the starting frame image, where M is less than or equal to N;
[0159] Determine the third element from the target frame image;
[0160] Based on the image content of the third element in the M-frame image, identify the feature changes of the third element, which include at least one of position change, pose change, or state change.
[0161] Based on the feature changes of the third element, generate a feature change description text;
[0162] Store the target image frame and the text describing the feature changes.
[0163] In this way, compared to storing all pixel information of the M-frame images, storing the target image frame and the text describing the feature changes can effectively reduce the storage space occupied by the M-frame images, and further achieve the goal of saving multimedia resource storage space.
[0164] It should be understood that, in this embodiment, the input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0165] The memory 709 can be used to store software programs and various data. The memory 709 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 709 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 709 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0166] Processor 710 may include one or more processing units; optionally, processor 710 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 710.
[0167] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described multimedia resource storage method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0168] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0169] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described multimedia resource storage method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0170] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0171] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes of the multimedia resource storage method embodiments described above, and can achieve the same technical effects. To avoid repetition, it will not be described again here.
[0172] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0174] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method of storing multimedia resources, characterized by, Applied to electronic devices, the method includes: Determine the target frame image from the multimedia resources, and M frames of images with the target frame image as the starting frame image, where M is an integer greater than 1; Determine the third element from the target frame image; Based on the image content of the third element in the M-frame image, identify the feature changes of the third element, wherein the feature changes include at least one of position change, pose change, or state change; Based on the feature changes of the third element, generate feature change description text; The target frame image and the feature change description text are stored; wherein, the method for storing the target frame image is as follows: The image content of the target frame image is divided to obtain a first element and a second element. The second element includes at least a portion of the image content of the target frame image, and the first element includes the image content of the target frame image other than the second element. Generate image description text based on the image content of the second element; For the target frame image, the image description text and the pixel information corresponding to the image content of the first element are stored.
2. The method of claim 1, wherein, The step of generating image description text based on the image content of the second element includes: Based on the image content of the second element, generate and display the initial image description text; Receive edit input for the initial description text of the image; In response to the edit input, image description text is obtained.
3. The method according to claim 1 or 2, characterized in that, After storing the image description text and the pixel information corresponding to the image content of the first element, the method further includes: Receive the first input; In response to the first input, the element content of the second element is generated based on the image description text; Based on the pixel information, the image content of the first element is obtained by parsing. A second image is generated based on the element content of the second element and the image content of the first element, wherein the image content of the second image corresponds to the image content of the target frame image; The second image is displayed.
4. The method of claim 1, wherein, The process of dividing the image content of the target frame image to obtain the first element and the second element includes: When displaying the target frame image, a second input to the target frame image is received; In response to the second input, a first element in the target frame image associated with the second input is determined; The image content in the target frame image other than the first element is determined to be the second element.
5. A multimedia resource storage device, characterized by include: The second determining module is used to determine a target frame image from the multimedia resources, and M frame images with the target frame image as the starting frame image, where M is an integer greater than 1; The third determining module is used to determine a third element from the target frame image; The recognition module is used to recognize feature changes of the third element based on the image content of the third element in the M-frame image, wherein the feature changes include at least one of position change, pose change, or state change; The fourth generation module is used to generate feature change description text based on the feature changes of the third element; The second storage module is used to store the target frame image and the feature change description text; wherein, the storage method of the target frame image is as follows: The image content of the target frame image is divided to obtain a first element and a second element. The second element includes at least a portion of the image content of the target frame image, and the first element includes the image content of the target frame image other than the second element. Generate image description text based on the image content of the second element; For the target frame image, the image description text and the pixel information corresponding to the image content of the first element are stored.
6. The apparatus of claim 5, wherein, The step of generating image description text based on the image content of the second element includes: Based on the image content of the second element, generate and display the initial image description text; Receive edit input for the initial description text of the image; In response to the edit input, image description text is obtained.
7. An electronic device, comprising: It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1-4.
8. A readable storage medium, characterized by, A program or instructions are stored on the readable storage medium, which, when executed by a processor, implement the steps of the method as described in any one of claims 1-4.
9. A computer program product stored in a storage medium, the program product being executed by at least one processor to implement the steps of the method as claimed in any one of claims 1-4.
Citation Information
Patent Citations
Image description generation method and device and electronic equipment
CN110717498A
Preview cover generation method and device, electronic equipment and storage medium
CN111880888A
Video data processing method and device, equipment and storage medium
CN117998039A